All posts
Math AI News June 17, 2026 4 min read 16

First Proof’s second batch tests whether AI can survive mathematical refereeing

Harvard’s latest First Proof report shifts attention from model claims to the human work required to check research-level solutions.

The referee is part of the experiment

The second First Proof batch asked AI systems to attempt ten previously unpublished research problems, then placed the proposed solutions in front of living mathematicians. The report’s most important contribution is methodological: it treats verification as a measurable part of the pipeline instead of assuming that a confident proof-shaped answer is self-authenticating.

Checking is not clerical work

A referee must reconstruct definitions, test edge cases, locate hidden dependencies, and decide whether an argument actually proves the stated claim. That labour can be harder than producing a plausible first draft. The finding matters for anyone building AI research tools: a system that accelerates conjecture generation may still create a new bottleneck if its outputs are expensive to audit.

A healthier research workflow

The practical model is collaborative. Let an AI propose routes, formalise routine steps, or search related literature; ask humans to choose the question, challenge the assumptions, and certify the result. Evaluation should report both the model’s success rate and the time expert reviewers spent reaching a verdict.

Read the report