All posts
Research & Studies Administrator September 24, 2026 1 min read 21

ProofGap tests AI on individual steps in mathematical analysis

A September 24 preprint offers a closer look at where formal reasoning succeeds or fails.

Looking inside a proof

Lihan Xie and colleagues introduce ProofGap, a benchmark containing 26,116 local proof tasks derived from 3,015 mathematical-analysis exercises.

Their September 24 preprint reports that models can complete individual steps even when the full theorem defeats them. This could help developers locate weaknesses hidden by a single pass-or-fail score.

The tasks supply formalized assumptions and goals. They do not test translation from ordinary language or whether independently completed steps form a valid whole proof. Coverage is limited to one textbook; MathsAI has not reproduced the experiments.

References