logoalt Hacker News

Sharing AI progress in mathematics

455 points • by OfficialTurkey • yesterday at 10:17 PM • 383 comments • view on HN

https://github.com/openai/math

https://github.com/openai/math/tree/main/preprints


Comments

zone411 • yesterday at 11:16 PM

A quick check shows that this list claims to fully solve 90 of the top 500 open problems in math (https://proofatlas.ai/open-problems/).

The highest ranked would be:

| 22 | Hilbert’s tenth problem over ℚ |

| 29 | Unique Games |

| 31 | Anderson-model extended states |

| 37 | Spacetime Penrose inequality |

| 48 | Nonexistence of Landau–Siegel zeros |

| 52 | Baum–Connes |

| 78 | Abundance |

| 80 | Hadwiger |

| 87 | Bose–Einstein condensation |

| 92 | Two-dimensional entanglement area law |

➕ show 3 replies
xanderlewis • yesterday at 11:36 PM

As Kevin Buzzard recently said:

> In a 2020 piece in the Notices of the AMS, I asked the following question: “If one human had an understanding of all of modern pure mathematics simultaneously, how much further would they immediately be able to see?” Six years later we are beginning to understand the answer to this question.

➕ show 3 replies
NotOscarWilde • yesterday at 11:24 PM

As a TCS/scheduling person, this one is definitely of lesser importance than UGC, but it has been an open problem since the book of Garey and Johnson in 1979:

A Polynomial-Time Algorithm for Three-Machine Unit-Job Scheduling [1]

Since some people talk about small numbers that pop up in integer multiplication results, here a completely different number appears:

Theorem 1.1. Let an explicitly listed finite directed acyclic graph specify the precedence constraints on n >= 1 nonpreemptive unit-length jobs on three identical machines. There is a uniform deterministic algorithm that constructs a feasible schedule of minimum makespan. Given also an integer deadline 1 <= T <= n, it decides feasibility exactly and returns a schedule whenever the answer is affirmative. Both tasks can be performed in O((L + 2)^150020) steps on a deterministic multitape Turing machine, where L is the total binary input length.

That is some crazy exponent -- plus an interestingly old computational model to boot; not something that is natural to most of us. I have no capacity to check its correctness today, but I hope it is true purely for the exponent.

[1]: https://github.com/openai/math/blob/main/preprints/A-polynom...

prideout • yesterday at 10:54 PM

This includes a proof of Barnette's Conjecture, which is one of the graph theory conjectures that I tried attacking with SOTA models a few months ago. I like it because it is easy to understand with a basic knowledge of graph theory. I spent quite a bit of time on it and failed. Their proof looks approachable at first glance.

https://github.com/openai/math/blob/main/preprints/Paired-st...

➕ show 1 reply
enoether • yesterday at 10:34 PM

Unique Games Conjecture [0] is a seminal conjecture in Complexity Theory, and is an underlying assumption for many, many inapproximability results. A valid proof is a big deal!

[0] https://en.wikipedia.org/wiki/Unique_games_conjecture [1] https://github.com/openai/math/blob/main/preprints/The-Uniqu...

➕ show 4 replies
schleck8 • today at 12:19 AM

Levent Alpöge (Anthropic mathematician) comment on the significance:

> Sure, mathematical history features a lot of incredible developments, like the invention of proof, zero, or the computer, and on the great problems our progress has been over timelines measured in decades or centuries. Obviously this technology didn’t appear today, but blurring our eyes a bit to combine the past ten years, with today a measurement of those developments, there is nothing comparable.

➕ show 2 replies
gizmodo59 • yesterday at 10:34 PM

This is significant progress and released without all the drama. Some very important progress in Reinmann, Hodge and unique games theorem. Point the repo to your agent and ask for the significance! In a way this is probably 50-100 years of math progress by humans

➕ show 3 replies
bcatanzaro • today at 1:10 AM

“I think at the heart of this issue is that humans have two competing natures: a tendency to compete and a capacity to appreciate beauty,” said Kai Shaikh, a graduate student in mathematics at the University of Toronto. “To me this seems to be a case of the former attempting to strangle the latter.” [1]

Beauty can be appreciated even when it is vast, even when it is beyond one's comprehension. I don't think this release should be primarily viewed as an outcome of competition. Instead it is revealing truths about the universe that were always there and always beautiful, even if we hadn't seen them yet. I believe there are infinitely more such beautiful truths currently hidden and waiting for us to discover.

[1] https://www.nytimes.com/2026/10/06/science/openai-math-probl...

sebmellen • yesterday at 10:33 PM

It’s fascinating to read through the reasoning traces: https://github.com/openai/math/tree/main/reasoning_traces

Look at one of their examples of an initial prompt: https://github.com/openai/math/blob/main/reasoning_traces/re...

➕ show 2 replies
againstapples • yesterday at 11:31 PM

As an AI "doomer" can I ask the non-doomer people here how you interpret the significance of results like these, and what kind of progress you expect to see in the next 1-5 years?

Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?

➕ show 14 replies
foota • yesterday at 11:14 PM

From their github: "The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking." That's pretty crazy.

davegoldblatt • today at 1:46 AM

Verified Riemann Zeta in Lean: https://github.com/davegoldblatt/openai-zeta-proof-check

➕ show 1 reply
kingstnap • yesterday at 10:54 PM

Some of these are interesting ngl.

109. Integer multiplication below n log n

Surprising that this is possible.

158. The Euclidean plane cannot be colored with five colors.

Only 6 and 7 remain!

376. Universal computation in forced Navier–Stokes flows.

Morning coffee proven turing complete

➕ show 2 replies
7373737373 • today at 1:33 AM

It may be useful to publish a formalization of ALL known mathematics at this point. Like every book ever printed, every paper on arXiv etc.

How many Gigabytes would that be, compressed? Wikipedia once fit on a DVD

This might also allow for some interesting meta-mathematics

➕ show 1 reply
dekhn • yesterday at 10:49 PM

I'm a software engineering/biology/ML guy who loves when clever math ideas get turned into real solutions (https://en.wikipedia.org/wiki/Compressed_sensing). I am curious if any of the results have immediate applications in any kind of engineering or science.

It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.

➕ show 1 reply
patcon • today at 2:02 AM

I'm a little concerned, but I'm also glad they're publishing quick before the department of war starts making national security claims of related to using the knowledge for cryptography and/or weapons

karahime • yesterday at 10:27 PM

Extremely unfortunate that gate keeping got to the point where they felt the need to ask for permission to share math.

➕ show 4 replies
open592 • yesterday at 10:41 PM

Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?

Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?

➕ show 15 replies
pavitheran • yesterday at 10:46 PM

From the GitHub description: “On average, each result used 3 hours of ChatGPT Pro thinking compute”

➕ show 4 replies
binlog • yesterday at 10:31 PM

So happy this is shared on GitHub rather than some gatekeeping paid journal. Truly a new age for science.

➕ show 3 replies
ks2048 • yesterday at 10:44 PM

I think they should put human names on the papers as someone who has reviewed the result, even if just a preliminary review. (I’m assuming they didn’t just pipe their model output directly to the internet and these had some amount of review?)

➕ show 4 replies
trostaft • today at 1:14 AM

Not much computational mathematics here, but I do see some of interest in the MCMC community. In particular, 139, 93, and 101 are very interesting. Also (I'm not too familiar), I remember attending some talks on the Crouzeix conjecture 325 attacks, should sharpen some rates for Krylov methods. Obviously need to read deeper, some of these don't have corresponding Lean formalizations.

Cool!

TheMrZZ • yesterday at 11:45 PM

These results are wild. Several individual findings are crazy good and use mostly unexplored methods (the improvement over Riemann for example)... I'm pretty sure some of these results would have been Fields-worthy.

But having so many of them at once? Damn. We really live in the future.

➕ show 1 reply
dualvariable • today at 1:59 AM

How many of these results are incorrect?

I doubt the answer to this is "none".

And how many of them are just exploiting some loophole that will need to be closed in the problem definition?

AmazingEveryDay • today at 1:50 AM

I think Alan Turing would be quite intrigued by these developments, and maybe wondering what took so long.

sigbottle • yesterday at 11:42 PM

Unique games conjecture and matmul <= 2.25. What the hell.

➕ show 3 replies
karannb • today at 1:07 AM

I think such "process-oriented" fields shouldn't be subjected to mass automation (or at least not till humans have sufficiently leveled up our game and understanding). What I mean is progress in math comes from having gained a deeper understanding of the problem for subseq

Xcelerate • today at 1:52 AM

> 241. Rigidity of the Turing degrees. Every order automorphism of the Turing degrees is the identity.

Wow. This is just crazy.

chickenjoseph • today at 1:50 AM

This achievement feels like a reasonable candidate for the moment where LLMs are officially "more intelligent" than any person. How can we justify moving the goalposts yet again? How can people have grown so numb to seeing advancements that they don't read this as significant? I feel like I'm standing at the foot of the exponential. I am not excited for the future, and I don't see how humans retain meaningful control over the future if we continue on this trajectory.

I have been lurking for quite some time. I made an account to post this, but I honestly don't know what to say. I would like to get off this wild ride.

➕ show 1 reply
bashtoni • today at 1:05 AM

Is AI going to put Mathematicians out of a job, or is it going to create many new jobs reviewing proofs it creates?

I'm not sure it's clear right now.

➕ show 1 reply
rinconrex • yesterday at 11:59 PM

The math equivalent of AI code reviews piling up. I wonder where the incentives will align and the equilibrium turns out.

rifty • today at 1:40 AM

As AI works steadily through the already discovered unsolved problems in mathematics, who is currently discovering new ones?

karannb • today at 1:11 AM

I think such "process-oriented" fields shouldn't be subjected to mass automation (at least not till we as humans have sufficiently leveled up our game and understanding).

What I mean is progress in math comes from having gained a deeper understanding of the problem for subsequent attack of more problems and IMPORTANTLY applications! Right now the first one is trivially satisfied (given oai maintains some memory across models) but the second one is not! It's generating proofs faster than anyone can validate and so only the model can use these. Consequently if it just keeps doing more theory it's... not very helpful or at least not optimally helpful. This is just bragging rights for now.

More "application-oriented" research would be awesome, where it tries to achieve some desirable effect and then produces relevant theory and experiments around it. Fields like CS, Physics, Chemistry, etc. This would also benefit a wider section of the population rather than the 10 people who understand most of these proofs.

avd201 • yesterday at 11:57 PM

Wow, FFT faster than O(nlog(n))? I wonder if that will open the floodgates for further improvement or not. I don't understand anything about most of the fields these results touch, but I can say that this in particular is very surprising.

➕ show 1 reply
baggy_trough • today at 2:16 AM

Stochastic parrot truthers in shambles.

closetheloopdev • yesterday at 11:54 PM

Hopefully the techniques and results here will be in the training dataset for the next models, so that each new release will give us more interesting techniques and results!

It seems that OpenAI has a proof machine that keeps multiplying fruitful proofs!

TeeWEE • today at 1:24 AM

This feels like a huge AI slop dump. Who validated these results. Why are they not mentioned. Human understanding is key here.

In my experience AI (frontier models) sometimes does weird stuff that needs human review. Not that’s incorrect but sometimes overly complex language or weird use of language.

curtis-jm • yesterday at 11:27 PM

You can read the papers here: https://hub.valency.io/collections/openai-math

lf88 • today at 12:12 AM

In some ways, this feels more like an ominous warning about the times to come than something to celebrate.

NegativeLatency • yesterday at 11:51 PM

Why should I care?

➕ show 1 reply
yewenjie • yesterday at 10:57 PM

A lot of these seem to be proving conjectures rather than finding counterexamples, a lot of people used that to claim that these models are not really smart/creative etc.

That copium didn't last for what, three months?

➕ show 1 reply
xydac • yesterday at 11:56 PM

i wonder what it means for maths researchers, and how it aligns with how they approach math problems.

➕ show 1 reply
i_idiot • today at 12:48 AM

If only AI can better humans in meditation...

lokl • yesterday at 11:58 PM

Do applied math next.

matapassiones • today at 12:00 AM

Valency has the papers up on Valency Hub

➕ show 1 reply
aaraujo002 • yesterday at 10:44 PM

The Advisory Group states in its recommendations [1]:

"We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."

To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too?

[1] https://agmai.org/general-sep29/

➕ show 10 replies
connor11528 • yesterday at 11:07 PM

will this make the math for building data centers work?

🔗 View 23 more comments