I can't relate to this at all. AI models will surely get better at writing "enjoyable proofs," but for now the situation is what it is. You're passionate about this problem, right? But you don't want to do the work to understand the result? Fine. There's a new generation of younger, hungry mathematicians that are highly interested in figuring out why the result is true and I am sure they'd be happy to wade through it and spoon-feed you the answer instead. Maybe they should be running things.
Very sensible comments. It is along the lines of the fury I get when I am confronted with an 11 page dump of an issue analysis created by an AI agent that makes no sense but I have to go through because customer shared it.
If you did’t bother to write it, I shouldn’t be bothered to read it.
Perhaps AI agents can have their own publications and magazines where they are the chairs and associate editors and reviewers.
A lot of the comments are claiming that "no one is forcing them to engage with AI proofs" and that's not the case, as explained in the article. The author is forced to engage with the public by the very nature of being a prominent researcher on this problem. The public is drowning him in messages regarding this result. So yes, he is being forced.
Academics have always been required to engage with hacks and cranks to some extent; the deluge of AI proof writing has only exacerbated the problem.
"not at the Lean code, since I know very little of the actual usage of Lean, and that code was enormous"
This part I don't understand. Not that anyone should read the entire Lean code of any proof, but if the statement of the theorem to be proven in lean seems to be correct, then I would think there would be at least some interest if in fact there was a formal proof (which might or might not correspond to the written proof) of something I was working on. That to me would be interesting. Or you are saying you doubt the validity of the formal proof, which would also be interesting. But saying it is of no consequence doesn't make any sense to me.
I agree, and I disagree. There are plenty of published mathematical papers that are just as poorly written as OpenAI's. Nobody says nothing because the authors are big names. In some cases, the proofs are not even correct, but everybody has a feeling the result are true nonetheless, so they pretend not to see it. So I agree that OpenAI should have done a better job of writing down the results, probably by paying working mathematicians like Anthropic did. But I disagree that this low-quality writing is somehow a good reason to be angry at OpenAI specifically, otherwise you would have to be angry at a lot of people.
The response by some in the field of mathematics to this repo is ... I guess not unexpected; but it's quite disappointing.
I sympathize with those who've worked on some problem for years and now don't have something to work on; it's been a part of their identity. I also especially sympathize with those whose career tracks and plans were thrown in disarray.
That being said, I absolutely cannot understand how one can't be excited and happy and enthused about these advances in one's field. Assuming just that the ones with formal lean proofs are actually true, these are reportedly huge advances. Even if folks don't understand it YET.
This is the key bit.
> But, back to the Partition Principle. I took a brief look at the preprint released by OpenAI (not at the Lean code, since I know very little of the actual usage of Lean, and that code was enormous). It sucked. It is unclear, muddled, and has a strange structure.
Incredible lack of curiosity. The Lean artifact shows that there is a proof. Maybe the natural language writeup sucks (maybe it doesn't even correspond to the Lean proof!) but the proof is there and if he were really interested in the problem he would try to understand it. Rather, his revealed preference is that what he's really interested in is good style in academic papers.
At first I thought this was going to be more Luddite babble, but it makes a good point. OpenAI isn't contributing if they are make unreadable papers. They should use a little more of their compute to nail interpretability. The difficulty will only get worse as AI plow deeper into the frontier and produce increasingly alien looking output. I suspect it's a workflow issue. If not, it's a bad oversight if the current generation of models are capable of making mathematical breakthroughs but can't explain how they build on existing frameworks.
Maybe the reason the AI could find this proof is exactly what the author is complaining about: that it left the beaten path of theorems expected in a paper like this and went off in an unexpected direction.
I'm just amazed at how quickly we moved from "AI is just a parrot and can't do anything useful" to "AI can solve toy problems but not anything of value" to "when AI pushes the state of the art, it can't quite get the proper citations in its preprints".
In short, we’ve reinvented cranks sending unsolicited, poorly written putative proofs
Mathematicians struggle with the same problem as software engineers; you can let AI generate the artifact, but to understand fully what is going on is challenging. Perhaps even more for mathematicians.
Do you rely on the tests/Lean to accept correctness or not…
I would have expected the math community to celebrate all of this since they got into math for the love for mathematics rather than the love for tenure. Their own world will not end - more mathematics will require more mathematicians to make sense of it all. And if all of it leads to some form of superabundance they are looking at the best possible future: they’ll be able to do math forever without having to worry to get paid for it.
>OpenAI drops some hundreds of "solutions", incomprehensibly written "solutions", and we are all expected to jump on them and what? Appreciate their contributions? Why?
...
>So, no, I will not be sending Sam Altman a bottle of whisky anytime soon, nor I am planning on spending my time reading through that paper and trying to make sense of it.
Think about a hypothetical circumstance where we get radio communication with some aliens on another planet. They send over tons of math to help us advance our tech, we know the math they're sending us is correct, but their explanations are really hard to work through because they aren't humans and the math is so different from anything we've done. Should we whine about the results they sent to us and refuse to engage with it?
As a whole, the math community would want to simply ignore the distractions from AI companies … they are using math problems as medium to promote their commercial interests.
> not at the Lean code, since I know very little of the actual usage of Lean, and that code was enormous
Dismissing results on the basis that Lean code is too long disqualifies this opinion. It is not hard at all to read the Lean result statement, even with very superficial Lean knowledge.
There is no nice way to tell someone that you’ve scooped them, and this is industrial scale scooping.
A few papers have been retracted, but it looks like many are withstanding intense scrutiny. Lean is making the results more likely to be correct, but I think making them harder to understand.
The world has changed and you’ll know a math department is making a serious attempt to adapt when it teaches a required Lean course in freshman year.
> "why you should be very angry at OpenAI"
Please don't tell me how I should feel. Stick with the facts.
Love this!
> This is why I generally avoid using AI for mathematics (I am happy to ask LLMs to consolidate information for me, or to generate a useful infographic, or to proof read an email, etc.)
In other words, the author is OK with using LLMs to replace data analysts (that could consolidate information), to replace graphic designers (that could generate infographics), and to replace editors (that could proofread an email). But don't you dare use LLMs in their mathematics.
openai will take expert responses like this and improve the next set of papers
it won't be long before there's no more low hanging fruit like this to complain about, and the writing / explanations of the results are superhuman as well
separately, i really liked the author's denial-of-service analogy. super useful practical framing
>That's not what you'd expect from a serious preprint claiming to solve a problem.
...
>It seems to me, that the AI tech companies would like us to conform to their standards, rather than spend the time and energy to conform to ours.
The entire argument is ugly and unfamiliar methods are used, and nobody's got time to drop everything and understand all that, meanwhile OpenAI didn't even ask us about this, therefore we should reject the paper as "unreadable" and you should be "very angry at OpenAI" today.
No, it's the press where you need to direct your anger. That is the nature of the beast. And being arrogant at the time of mass-disruption is a great way to lose control of everything. Maybe Sam should be sending the bottle of whisky to you.
Very well written summary of the current situation
Imagine being someone who is working on one of these problems. You have no good guarantee that the problem was solved, but you will have the horrible homework of reading the AI slop. Also, if you do have something interesting to say about the problem, people will have less enthusiasm about it now
If he doesn’t like the paper’s structure he can just prompt gpt pro to review and amend.
I know the model that produced these proofs is still private, but it’s worth a shot tackling the proofs with the current consumer-available frontier.
Great article. My take on it is :-
"Er, can you check this for free? ... It would be great for my share price if you could, would really give the investors a badly needed shot of confidence!"
Here’s an idea I would love to see play out. Have one of the labs train a new model, using cutting edge architecture, on a whole lot of data and math papers from before 1905. Then see if it can come up with, or even understand, Einsteins theory of relativity.
If there is a formal proof, and it is the proof of your precise statement (which I imagine is easy to check, otherwise I think that mathematicians would not accept the Navier Stokes result so quickly), then there is no way to ignore the result, however badly it is written.
This was always the essence of mathematics, and it will stay this way whichever statement by whomever is made.
As much as I personally despise altmans, "darios", and their bootlickers, this is one aspect which is undoubtedly "good for the mathematical community" as a whole. The fact that the validity of your statement does not depend any more on an expert opinion of some person with grants, but as it always should have had been, just on the validity of the chain of deductions.
It certainly feels that we are witnessing the early days of LLMs and math, similar to the early days of LLMs and code. I feel this entire post sounds so similar to angry computer engineers posts from 1 or 2 years ago. If we continue on this path, eventually many of these posts will ripen like milk in a desert.
It’s fine to dislike AI slop and not engage with it. But I would claim there is a difference between a human submitting a sloppy paper to a journal vs producing one with AI. The former is lazy and unprofessional, the latter is an interesting experiment. I appreciate not wanting to engage in an experiment you didn’t sign up for, but therein lies the difference between this with an attitude of embracing new technology and those wanting to stick to the old. I think Terrence Tao has had a very interesting attitude to AI recently and had also had some fruitful outcomes from it.
I would admit seeing more mathematicians irritated is a good popcorn show.
For what's worth it, this batch of results are not very tight intentionally by OpenAI and promising mathematicians already started to consume them and improve the results, while someone is still complaining on it.
> OpenAI drops some hundreds of "solutions", incomprehensibly written "solutions", and we are all expected to jump on them and what? Appreciate their contributions? Why? I am trying to finish several papers, I am supervising a number of Ph.D. students, and I have a lot of active research of my own to do. When am I supposed to sift through a badly written paper? Why should I bother, when they don't bother to communicate better?
Excellent point.
I don't blame the AI for this - I blame OpenAI.
Literal slop grenade (see https://fortune.com/2026/09/17/shopify-tobias-lutke-ai-slop-...)
The mathematicians who don't approve of the deliverables should boycott the proofs, that is the only way OpenAI will get what they deserve on this one.
I see an argument about the paper being poorly written, which I believe, but I don’t see how that relates to honest final paragraph about mathematicians being chefs or whatever.
It’s clear from the author’s tone about lean, emails, infographics, etc. that he thinks automating those away is fine. Why should math be any different?
Is a proof that cannot be understood worthless? How would this be framed philosophically?
In fact, the academic system is a kind of worldview created by humans. And as it is shared and the community grows, the problem will gradually become more complex. Because when a discipline develops sufficiently, just as in a mine where rich veins are easy to extract early on but become very hard to extract once much has been dug out... in that sense, as things gradually become more complex, once a certain threshold is reached, won't scholarship surpass the limits of human understanding? Of course, scholarship is entirely for humans, but at some point the system itself may face its limits, and then wouldn't it again reduce the existing normalized minimum within that discipline and establish a new normalization of a new logical system?
In my view, perhaps for very complex work like today, AI will do it, and then there will be work that normalizes and further simplifies the results of that AI. Then, coming back to the human fold, if humans create the initial skeleton, the LLM will learn that again and it will become complex work again, and won't this create a continuing cycle?
I think verification and understanding can be separated. If the proof targets a correctly formalized proposition and passes a reliable proof checker, isn't it valuable? We have obtained knowledge justified as true, but there is simply no new theory that understands that knowledge. As was the case with the Four Color Theorem...
I am always curious what shape the newly compressed new discipline will take. At that time, I hope even people like me, who are intellectually behind, will be able to learn that discipline.
I've seen my entire profession vanish overnight due to AI... well, not vanish. But, yeah, software development is WAYYYY different. And I couldn't be happier. I think it's amazing. I see the productivity boost. Even if it means I can't add nearly as much value as I used to.
Sorry, but I don't get why mathematicians are so upset. Like, just accept the knowledge and insights and acceleration in your field! If it isn't "fit for human consumption" because an AI produced, okay... it soon will be explained ELI5 by even better models.
> OpenAI drops some hundreds of "solutions", incomprehensibly written "solutions", and we are all expected to jump on them and what? Appreciate their contributions?
This is a strawman. OpenAI didn't say they are expecting all mathematicians to read the solutions, incomprehensible or not.
So mathematicians are upset with OpenAI for solving "their" math problems. Software engineers are even more affected by AI, yet mathematicians seem to be reacting more strongly. I don't get why.
> Indeed, at least two people asked if I plan on sending Sam Altman a bottle of whisky, as promised in my Problems page. The answer to that is no. And let me explain to you why
> I took a brief look at the preprint released by OpenAI. It sucked. It is unclear, muddled, and has a strange structure.
https://img.getfn.io/images/e3245d47a159592b06570ffbd64d5af8...
If OpenAI actually spent time writing good proofs that are readable, mathematicians would show *even more outrage*. So let’s not pretend this has anything to do with readability. It has all to do with people protecting their status in society.
> if this was an academic paper submitted to a journal, it should be issued a desk rejection for the quality
If Mr Tao had submitted a shitty, poorly-written proof of a famous outstanding problem, no journal would reject it. That extends to anyone with sufficient credibility. They might ask him to keep at it and fix it up, but nobody would begrudge him putting his shitty (but ultimately correct) draft of arXiv while he did so.
We're all getting disrupted, we all have feelings about it, but from the perspective of a software engineer who's been dealing with all of this for several years now, this post is just cope.
[dead]
Wow talk about sour grapes! No one is forcing you or any other mathematician at gun point to engage with this release at all. Ridiculous drivel.
I just read a childish rant
This whole "oh no this problem is solved now who would ever want to work on it" thing is quite funny. Mathematicians, let me introduce you to something called bike shedding. Folks proved 1s and 0s are turing complete decades ago and yet we have 27 new JavaScript web frameworks every week (or every day or minute now with ai). Y'all will be fine. Thinking a solution is ugly and that you could do better is also a perfectly good motivator and most of the sciences and engineering are "ok, so we know xyz to be true about the world because like, I'm looking at it, but wtf is going on." The theorems were true false or otherwise before some random openai model solved them. And if you don't understand the proof nothing of significance has changed except you've got a bit of a hint now.
Cathedral and Bazaar. Maths professors are used to working diligently behind closed doors before releasing artifacts of high quality whose authorship they guard jealously. The researchers at OpenAI, who come from a software background, are used to working in the open, releasing anything to anyone and expecting nothing but also guaranteeing nothing.
The amount of time the OP attacks OpenAI for errors of form and not substance is unfortunate. Do they also attack amateurs who try to contribute like this?
I think it can both true that 1) OpenAI is being inconsiderate/harmful/<pick-whatever-adjective> with their math releases, and 2) there is now a treasure trove of mathematical results ready for the taking.
Yes, the situation sucks overall and mathematics as a whole is in a turbulent time now.
But it also sucks when mathematicians, who are considered experts on a particular problem, refuse to engage with breakthrough results about that problem. #2 above is still true regardless of where it came from or how hard it can be to absorb.
While reading "The Mathocalypse" post [0] by Scott Aaronson, Scott described his wife Dana's reaction to one of the newly solved results in her primary domain of expertise, on which she'd been working for decades.
After her initial shock, and annoyance with the format/style, she decided to start using Astra - for the first time - to help her understand the new result. And he reported in the comments that she had made a lot of progress understanding it in one day, and may be even excited to give a talk about it!
That seems like a much healthier attitude towards these new results.
Yes, everything else sucks about this messy period. But there are still diamonds (in the rough) in this drop that perhaps should be looked into. If the author is too busy, perhaps one of their students can take a look? Someone will, eventually.
[0] https://scottaaronson.blog/?p=10169