Sol Didn't Just Solve a 79-Year-Old Math Problem — It Rewrote Half of Mathlib in 3 Weeks
If you haven't been paying attention to what OpenAI's Sol has been up to behind the paywall, you're about to feel very small. Two months ago, ChatGPT dropped a bombshell: it had disproved Paul Erdős's Unit Distance conjecture, a problem in discrete geometry that had been kicking around since 1946. That alone would have been a headline. But what happened next is where the story goes from "impressive" to "genuinely unsettling."
Because Sol didn't just find the counterexample. It then proceeded to write the mathematical equivalent of a doctoral thesis in Lean — the interactive theorem prover — in three weeks. The output: 1.2 million lines of formalized, machine-checked Lean code. To put that in perspective, the entire mathlib library — the collective effort of hundreds of mathematicians over nine years — is 2.3 million lines. Sol did half of that, from scratch, in 21 days.
The Counterexample That Broke the Mold
Let's go back to May 20, 2026. That's the day ChatGPT announced it had disproved the Erdős Unit Distance conjecture using a deep result from 1960s number theory: the Golod–Shafarevich theorem. The proof was accompanied by testimonies from respected human mathematicians who had been given early access and had checked the argument. It seemed legitimate. But there was an elephant in the room: the proof hadn't been formalized.
Enter Kevin Buzzard, the Fields-adjacent mathematician behind the Xena Project and a Lean maintainer. Buzzard has spent nearly a decade arguing that interactive theorem provers are essential for the future of mathematics — specifically because human peer review keeps missing things. So his first question was obvious: "Is it formalized in Lean?" The answer was no.
That changed fast. Within a week, Fields Medalist Mike Freedman — now Chief Science Officer at Logical Intelligence, a company cofounded by Yann LeCun — emailed Buzzard saying their system had autoformalized the ChatGPT paper in Lean. But it was a partial formalization. It proved that the Golod–Shafarevich theorem implied the counterexample, but the 100+ pages of global class field theory underlying the theorem itself remained unformalized.
Three Weeks, 1.2 Million Lines
Then came June 26. Boris Alexeev, a researcher at OpenAI, announced on the Lean Zulip that he had steered ChatGPT's new model — Sol — to a complete formalization of the entire Erdős counterexample, assuming nothing beyond the axioms of mathematics. The code was public. Buzzard ran it in a sandbox and watched Sol prove nontrivial theorems about the cohomology of number fields. On his machine. From AI-generated Lean code.
- 1.2 million lines of Lean code generated in 21 days
- Complete formalization of the Erdős counterexample including global class field theory
- Proofs of hard theorems in the cohomology of number fields — self-contained from ZFC axioms
- All public — the code is out there for anyone to inspect, run, and build on
Buzzard himself wrote: "Perhaps it was at this point that the penny really dropped for me — large AI-generated developments of mathematics are inevitable." Coming from a mathematician who has spent the last nine years building mathlib by hand with a global community, that sentence lands like a hammer.
The Formalizing Fermat Workshop: A Glimpse of the New Normal
If you think this was an isolated experiment, the timeline gets wilder. In early July, Buzzard ran a five-day "Formalizing Fermat" workshop. Logos Research (one of several AI startups now specializing in autoformalization) sponsored it. OpenAI offered all 25 attendees free ChatGPT Pro access. Anthropic offered Claude Max subscriptions. The timing was deliberate: Sol launched publicly on July 9, right in the middle of the workshop.
- Attendees had access to Sol (OpenAI), Fable (Anthropic), and Logos' autoformalization tool simultaneously
- The workshop was a race to formalize Fermat's Last Theorem in Lean
- Multiple AI companies competing to provide the best autoformalization engine — in real time, in front of mathematicians
This wasn't a demo. This was the new reality of mathematical research: human mathematicians, AI models, and Lean all working in the same room, at the same terminal, on the same problem. The competitive dynamic between AI vendors added an edge that Buzzard described as "unprecedented."
What This Actually Means
Let's be clear about what Sol achieved. This wasn't a parlor trick or a benchmark cherry-pick. The Erdős Unit Distance conjecture is a real, open problem that had resisted solution for 79 years. Finding the counterexample required deep insight into number theory. Formalizing it required translating that insight into the brutally precise language of a theorem prover — where every assumption must be stated, every lemma proved, every edge case handled.
The fact that Sol could do this means we've crossed a threshold. The question is no longer "can AI do mathematics?" It's "how do we structure mathematical research when an AI can generate half of mathlib in three weeks?"
Answers aren't coming. But the code is public, the proof is checked, and the field of mathematics just got a lot more interesting.
Comments