I see two assumptions in the core argument, both of which are open to challenge.
1. AI can do proofs, but deciding which problems to solve, which math is useful, is by humans.
2. The math that's picked needs to be understandable by humans.
For #1, AI may be able to play a significant, if not a takeover, role for even figuring out what math is useful.
For #2, understandability by humans may be good for now, but could also turn out to be a significant constraint. Correctness is a goal, trust is an important requirement, human understandability may be an intermediary for that, but not necessarily the end goal.
In other words, the article may stand the current state of the art, but may not stand merely a couple years down the road.
Some arguments are based on "past experience..."
However, such experiences are not absolute truths and cannot be equated with the current situation.
Read by the author https://www.youtube.com/live/gPrWX8i1htM (with multiple mentions of the Wolfram language and such)
The desperately needed TL;DR is that perhaps the actual mathematics itself (i.e. proofs, calculations, etc.) can be done by AI, but why we do it and deciding which problems to solve can only be done by AI. Therefore mathematician do maths.
I'm not a mathematician, but this seems like a weak and slightly bizarre argument.
Wolfram gives a very important and sober take on the present state and future of pure math. I came away with the following key takeaways:
1. An essential goal of mathematics is human understanding. The computation of proof terms doesn't necessarily enrich human understanding. The proof of the four color theorem result is a good example, and formal verification/SAT solving gives many more: these are results that can be trusted up to our trust in the system used to produce them, and they can be used in practice, but they don't necessarily enrich our understanding. Imagine a computer with near infinite proof search powers set loose with the current human definitions, theorems, and understanding of mathematics. Suppose it constructs a proof for a new theorem at our mathematical frontier. The shortest such proof in terms of currently understood definitions and concepts could be so long and mechanical that the entire lineage of humans until the end of the universe could not finish reading it. So though it overlaps with the activity of mathematicians, this type of computational proof search is not mathematics as such. This is an important distinction that many people do not seem to grasp and some dismiss as cope.
2. The human activity of theory building, rendering otherwise monstrous proofs like the one I discussed above into light conceptual arguments a person can understand, appears at this time out of reach of models. Maybe they will do this in the future, but it is not yet the case. Human theory building drastically compresses the spaces of theorems and their proofs: this is why great theory builders like Groethendieck are so important to the field; grinding has its value too, but runs up against computational limits in both humans and computers. These limits are collapsed by the conceptual shortcuts created by theory builders.
Gowers has a nice and arguably better-grounded article on the mathematical capabilities of recent LLMs that I think is enlightening to read alongside Wolfram's bird's eye view of the implications of those capabilities: https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-a...
> back in 1988, when we first introduced Mathematica, there was also some of the same kind of talk about math being taken over, and made pointless. Of course that’s not how it worked out at all.
Makes sense.
> For me, its greatest use in mathematical pursuits has been its ability in effect to thematically mine the knowledgebase of human mathematics. ... Modern AI is, first and foremost, a way of leveraging the existing corpus of human knowledge.
AI is more a database of knowledge (stolen knowledge but let's leave that discussion asside) than a thinking machine. You can query a compressed version of millions of books.
That is very useful. (Are we already StarTrek-communists?
> generating useful mathematics is a much more exacting activity than generating language.
This is something that most people forget. Generative AI is mostly LLMs, and they are chatbots not mathbots.
> It’s a frustrating feature of modern times that someone like me gets sent many AI-generated documents every day that have the “statistical texture” of math papers, but that one at least expects have a very low probability of being meaningfully correct
And here is the trick. A million monkeys with a million typewriters may write a Shakespeare masterpiece. But they would not be able to differentiate it from garbage text.
> So, yes, there’s every reason to expect a bright future—now with some additional help from AI—for that most rarefied of human pursuits: research in pure mathematics.
Happy to hear that.
Much more rejections due to more submissions, and AI will probably not help you (unless you solve a major open problem).
Even frontier models still struggle at proving small conjecturers despite all the hype about major breakthroughs. It really depends a lot on the prompt, the type of problem, among other factors. But AI does not suddenly make publishing in a journal easier, although it does make it easier to produce papers.
I like Wolfram and always enjoy reading his posts. An irreconcilable thing here is Wolfram clearly wants the Mathematica language to be central to the development of math as a field, but it's proprietary and so there is no guarantee it will survive the dissolution of the company if/when that happens. With Lean & others you can fairly safely assume that a particular version will be archived somewhere, and so a proof formalized for that version can be checked at any time. Not so with Mathematica. It's unfortunate, Mathematica is a cool language, but that's just the structure of incentives at this time. This is without getting into what a proof formalized in Mathematica would even mean, the relative maturity of the kernel and the possibility of there being multiple kernel implementations for cross-checking, and so on.
> I’ve had the experience quite a few times now: I try to autoformalize something, and an AI will tell me “I did it; look, the proof checks out!” But, actually, in some sense it cheated: instead of formalizing what I intended, it found a (sometimes very squirrely) way to interpret what I asked so that it could successfully prove it. It’s often very hard to tell, though, that this is what happened—not least because the formalized versions of things (say as expressed in popular proof assistant systems) tend to be very low level, very verbose and very hard for us humans to understand.
This has largely been my experience in programming, too. Often when given a broad objective within an existing code base, even the frontier models seem to stand up a half dozen tests proving compliance and then add 3 new branches solving for those specific tests alone. My best guess is that our code base is quite far out of the distribution (and not for just good reasons). The agents are thus reduced to these tactics rather than extending and refactoring well known abstractions.