All posts
ProofGap tests AI on individual steps in mathematical analysis
A September 24 preprint offers a closer look at where formal reasoning succeeds or fails.
Looking inside a proof
Lihan Xie and colleagues introduce ProofGap, a benchmark containing 26,116 local proof tasks derived from 3,015 mathematical-analysis exercises.
Their September 24 preprint reports that models can complete individual steps even when the full theorem defeats them. This could help developers locate weaknesses hidden by a single pass-or-fail score.
The tasks supply formalized assumptions and goals. They do not test translation from ordinary language or whether independently completed steps form a valid whole proof. Coverage is limited to one textbook; MathsAI has not reproduced the experiments.