PromptZone - AI Prompts, Guides and Tools for Builders

Farrah Saleh
Farrah Saleh

Posted on

Is AI Mining Open Math Problems Dry?

Terence Tao flagged on Mastodon that large language models are rapidly solving open math problems that once served as shared challenges for human researchers. The post triggered a Hacker News thread that reached 92 points and 64 comments.

The core concern is that these problems function as non-renewable resources. Once an AI produces a solution, the problem loses its value as an unsolved benchmark that multiple researchers could attack independently.

What the Discussion Describes

Tao notes that many open problems in number theory, combinatorics, and analysis are now being fed directly into frontier models. The models generate candidate proofs or counterexamples at scale.

Unlike human researchers who publish incrementally, models can clear multiple open questions in a single training cycle or inference run. No central registry tracks which problems have already been attempted by AI systems.

Scale of the Activity

Early testers report that models now handle problems previously considered graduate-level or early-career research targets. One commenter listed specific problems from the 2010s that now return complete proofs from current systems.

No public dataset yet quantifies how many previously open problems have been closed by AI since 2023. The HN thread treats this absence of tracking as the immediate practical gap.

Community Reactions on Hacker News

Commenters raised three recurring points:

  • Reproducibility concerns when model-generated proofs lack human-readable intermediate steps
  • Incentive shifts for mathematicians who may avoid problems already likely solved by AI
  • Potential value in creating new classes of problems designed to resist current model capabilities

Several users suggested maintaining a public ledger of AI-attempted problems, similar to existing formal verification repositories.

Long-Term Research Impact

The non-renewable framing implies that the stock of accessible open problems shrinks with each successful model run. Fields that rely on a steady supply of unsolved questions for training and evaluation face a structural change.

Researchers who previously used these problems as calibration points now need alternative benchmarks that models have not yet seen.

Practical Steps for Mathematicians

Groups can publish new problem lists with explicit instructions that models should not be trained or prompted on them until a later date. Some departments have begun timestamping problem statements on arXiv to establish priority.

Formal verification tools such as Lean can record which problems remain open and which have received machine-generated proofs that still require human checking.

Bottom line: The supply of untouched open math problems is finite and currently being drawn down without systematic tracking.

Future benchmark design will likely treat unsolved problems as a scarce, time-stamped resource rather than an unlimited public good.

Top comments (0)