As OpenAI announced hundreds of purported solutions to some of the world’s most difficult math problems this week, Frontier Labs said it consulted an advisory group of elite mathematicians to avoid the controversy that arose when one of its models solved a long-standing problem in the field.
However, OpenAI did not meet these criteria, especially in that mathematicians emphasized the need for humans to understand mathematical results. This is especially worrying after a new paper highlights the gap between natural language and formally expressed solutions to million-dollar problems ostensibly solved by OpenAI’s models.
The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), sponsored by Princeton University’s Institute for Advanced Study, is comprised of nine distinguished researchers from institutions around the world.
At the end of September, the organization released guidelines for Frontier Lab to solve math problems. “It is ultimately up to the mathematical community to assess how well our recommendations have been followed,” AGMAI said in a statement about the latest set of proofs.
But the organization’s first request was to stop testing advanced mathematical problems with its own models. OpenAI’s release clearly states that it uses unsolved research questions in mathematics to evaluate its own models.
The advisory group did not respond when TechCrunch asked for a more thorough evaluation of OpenAI’s latest proof release. The institute clearly followed some of its principles, including publishing results as soon as possible and including information about how the model reached its conclusions. But not all. Only 10 of the 719 manuscripts included publication of the model’s train of thought.
For papers that people don’t understand, mathematicians suggested formalizing the proofs, but only 42% of the proofs published by OpenAI did not go through this process.
After all, it is not yet clear whether OpenAI is “responsible for ensuring continued human understanding” in accordance with the AGMAI principles when publishing its proofs. AGMAI proposed that OpenAI should help fund the research of human mathematicians needed to make the lab’s solutions meaningful in a practical sense.
“Problems are being solved autonomously by AI teleprompters, with no concern for the broader discipline itself once the original goal is ‘solved’. They do not understand the AI output well enough to answer questions about the results, give lectures, or interact with other disciplines,” wrote Terence Tao, a prominent mathematician who has criticized OpenAI’s approach.
This problem is exemplified in a paper published this week by mathematicians from the University of Cambridge and King’s College, London, which calls into question how frontier laboratories are tackling these challenges.
When an AI model solves a mathematical problem, it first creates a “natural language” explanation and then attempts to express the result in a programming language called Lean. Lean is a programming language that theoretically checks the accuracy of proofs by compiling them as code.
However, there may be a problem with how the model translates the natural language proof into code. The paper documents at least two inconsistencies between the natural language proof and the lean code behind the solution provided by OpenAI to a problem derived from the Navier-Stokes equations that describe the complex behavior of fluids.
These contradictions do not necessarily disprove either solution, but they do raise the question of whether we can simply rely on models to formalize our own solutions without human involvement. This is one reason why AGMAI required OpenAI to “include machine-readable metadata that correlates natural language and formal artifacts,” but Frontier Labs did not do this in these releases.
“As highlighted in this paper, due to the phenomenon of mistranslation, OpenAI and
Other automatically formalized Lean proofs should prima facie not be trusted without the same peer review process.
“Other proofs are also subject to scrutiny,” the authors of the Lost in Translation paper conclude.
Mathematicians emphasize that when new results are discovered by humans, humans take responsibility for them and engage with the broader community through papers, lectures, and seminars. This process deepens your understanding of solutions, helps you find strategies you can use to solve other problems, and allows you to apply your new knowledge to real-world situations.
When a model is prompted to solve a difficult problem and spits out a solution, “humans can’t understand the model when it’s released, and the work starts now,” Melanie Wood, a professor of mathematics at Harvard University, told TechCrunch.
If you make a purchase through links in our articles, we may earn a small commission. This does not affect editorial independence.
