A new survey maps the AI-for-mathematics field—and its measurement traps
A June 2026 survey brings benchmarks, formal proof, multimodal reasoning, and reporting discipline into one research map.
The field needs a map, not just a leaderboard
A new arXiv survey reviews artificial intelligence for mathematical reasoning across grade-school arithmetic, competition problems, geometry, formal proving, multimodal tasks, and expert evaluation. Its value is organisational: readers can see how apparently similar claims may rely on very different data, scoring rules, and verification procedures.
The measurement traps
Pass@1, majority voting, verifier-assisted sampling, and human review answer different questions. A model that succeeds after many attempts is not being tested in the same way as one that produces a correct proof on its first try. Likewise, a formal proof certificate is stronger evidence for validity than a fluent explanation, but it does not automatically show that the theorem was meaningful or that the model understood the strategy.
A useful reference for builders
Anyone creating a math AI tool can use the survey as a checklist: document the dataset, prevent contamination where possible, separate discovery from verification, publish failure cases, and report the cost of repeated sampling. Better reporting is a technical advantage because it makes progress reproducible.