OpenAI’s Astra Model solved 10 open math problems: This is how we know it’s true

OpenAI’s Astra Model solved 10 open math problems: This is how we know it’s true

OpenAI is known to drop research bombs with no great announcement, and the newest one requires a closer look than just the small paper where it is published. The internal version of Astra, the next big model to be released by OpenAI, has come up with solutions to ten mathematical problems and problems in theoretical computer science which have been unsolved for at least a decade, and in many cases for much longer. The topics include high-dimensional geometry, coding theory, group theory, quantum complexity and lattice cryptography. The two last items in the list solve the remaining two problems from the famous problem list of Paul Erdős. One refutes a conjecture from operator algebras made in the seventies. So, what is the right answer? How can you trust the AI when it claims it solved a math problem?

Digit.in Survey
✅ Thank you for completing the survey!

Also read: Best tablets for designers in India in 2026: Five picks for illustrators

This is the story that is worth telling, and it is a lot more interesting than what meets the eye from the headline figure. OpenAI did not simply provide a list of the ten counterexamples and ask mathematicians to trust it because of the claim of mathematicians. All the ten counterexamples were made into a proof certificate using Lean prior to their being published. A proof certificate is a tool called a proof assistant which is basically a programming language capable of encoding mathematical arguments in a format that could be understood by the computer and could be checked logically.

That’s an essential distinction at this point in time. Big language models are well-known to make confident, fluent, and sometimes totally made-up assertions, which is a problem that research labs haven’t figured out how to overcome yet. An assertion from a big language model that it has proved something is basically meaningless. An assertion from a big language model that its proof has passed mechanical verification to align with formal logic is an assertion of an entirely different order. OpenAI clearly recognized that difference. “OpenAI is responsible for the accuracy of the manuscripts and the Lean formalizations. The mathematical proofs come from the system.” 

Also read: India’s real chip opportunity is supply chain, says Lam Research’s Rangesh Raghavan

One should note what lean verification does and doesn’t do. It verifies the logical integrity of the reasoning process assuming a set of premises is provided. It doesn’t mean anything regarding elegance and profundity of a proof, or that it can be called “illuminating” by a professional mathematician. Mathematicians have discussed the idea that a proof is just as much about the understanding of a statement as its verification for a long time. A machine verified proof avoids that discussion altogether.

That is precisely the issue that is of primary importance for the future credibility of AI-based science. Given that the applications of these increasingly powerful models will move on from solving toy problems to tackling ever-harder challenges in physics, chemistry, and engineering, a process of formal verification of some sort may well become the norm, not the exception. The expanding application of Lean to this end is no accident. It is a pattern. If the AI community wishes for its scientific statements to be taken seriously and not simply viewed as marketing slogans, formal verification will be the requirement. OpenAI appears to have got the memo. Whether the rest of the AI industry has followed suit is an open question.

Also read: This is the year for refurbished phones: Yug Bhatia says rising smartphone prices is changing India’s market

Vyom Ramani

Vyom Ramani

A journalist with a soft spot for tech, games, and things that go beep. While waiting for a delayed metro or rebooting his brain, you’ll find him solving Rubik’s Cubes, bingeing F1, or hunting for the next great snack. View Full Profile