All posts
Research & Studies Administrator September 19, 2026 1 min read 20

AI teams learn to combine partial solutions on mathematics tests

A September 19 preprint finds that learned teamwork improves results across five mathematics and physics benchmarks.

When AI agents work through a math problem together

A Stanford-led preprint posted on September 19 studies whether AI agents can improve mathematical reasoning by learning how to collaborate. Instead of assigning a fixed role to each model for every problem, the researchers let a three-model team develop reusable patterns for discussion, checking and combining partial answers.

The team learned its strategies from 15 AIME 2024 training problems, then used the same strategies on held-out questions and other contests. Across five mathematics and physics benchmarks, the authors report 66.7% average accuracy, compared with 48.8% for the strongest individual member and 58.7% for a compute-matched single-agent control. The team also exceeded a selector that could pick the correct answer whenever any member solved a problem independently.

That comparison suggests the exchanges sometimes produced a solution missing from the agents' separate attempts. One reported example combines an invariant proposed by one model with a counting correction from another. Yet the results concern selected benchmark questions and specific models; they do not establish reliable research-level theorem proving. The paper is a preprint, and MathsAI has not independently reproduced its experiments.

References