AI Signal 403
An unreleased Anthropic model made progress on one of math’s biggest unsolved problems
An unreleased Anthropic model advanced work on the Riemann hypothesis by testing many ideas and increasing the lower bound of cases where the hypothesis holds.
The experiment shows that a language model can orchestrate dozens of sub-agents to generate and validate mathematical ideas over a day and a half. It demonstrates that AI can contribute to lowering the bound on a long-standing open problem, which may influence how engineers approach automated theorem proving. The result also highlights ongoing debate about credit and responsibility when AI produces mathematical work.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
The model tested 650 ideas using 60 sub-agents, with two sub-agents producing the core mathematical insights.
The work was validated by Anthropic’s in-house mathematicians and formalized in the Lean proof assistant.
The outcome fuels discussion in the mathematics community about AI’s role in proof attribution and the future of human-led research.
THE READ
What the cluster adds up to.
The unreleased model explored the Riemann hypothesis by generating many candidate ideas via a swarm of sub-agents, leading to a new lower bound on cases where the hypothesis is true.
Achieving this required prompting a non-expert to set the goal, letting the model run for a day and a half, coordinating 60 sub-agents, testing 650 ideas, and consuming 31 million in total, plus subsequent validation by in-house mathematicians and formalization in Lean.
The model did not produce a complete proof; it only improved the bound, leaving the hypothesis unsolved for the remaining cases; the approach depends on human experts to confirm correctness and to translate informal arguments into a proof assistant.
The experiment shows that LLMs can act as a catalyst for mathematical exploration, but it also reignites debate about attribution and responsibility when AI contributes to formal proofs, indicating that adoption will need new workflows and governance.
Written by elseif from the cluster below · checked for specifics the sources never containedTHE CLUSTER
↗