logoalt Hacker News

OpenAI withdraws three mathematical results

242 points • by sashank_1509 • yesterday at 7:05 AM • 541 comments • view on HN

https://github.com/openai/math/blob/main/history.md


Comments

chubot • yesterday at 3:37 PM

So were the 3 withdrawn proofs the ones without Lean verification?

If so, why did they mix proofs that were verified with Lean, and proofs in natural language?

I was wondering that while reading Aaronson's blog:

https://scottaaronson.blog/?p=10169

Or at least, we’re pretty sure that it’s a proof! There’s a Lean certificate, as there are for some of the other 372 breakthrough results (not all of them). But it also appears that no human has understood just about any of these proofs yet

It seems that the obvious thing to do would be to release in TWO parts: the ones that are verified, and the ones that might have some good ideas but also might have some mistakes. Presumably the latter would be much more epxensive for humans to verify.

➕ show 1 reply
ijustlovemath • yesterday at 1:54 PM

I think that on closer inspection, a lot of these fully AI generated proofs will fall apart. Even in Lean, you can build theories which compile but nonetheless state something different than what you actually intend. It's just that the volume of proof is so staggeringly large that it will probably take years before we find the issues, a la abc conjecture

➕ show 6 replies
ThePhysicist • yesterday at 12:59 PM

I find the paper about beating O(n log n) for integer multiplication also quite fishy, not sure but it seems like too good to be true, I feel like there must be a subtle flaw in that. Maybe that's just me hating these small numbers in the paper, but it seems wrong, unnatural even! I would be similarly skeptical about a physics paper that claims to be able to exceed the speed of light by a tiny fraction. There's no reason n log n is the natural limit here but I see a few good intuitions so having something else that can't be represented in an elegant form seem very "unmathematical" to me.

➕ show 5 replies
renyicircle • yesterday at 11:25 AM

This is what it looks like when software engineering practices meet mathematics. "openai/math release 1.3.42: retracted papers 139 and 140, fixed a sign error in paper 47, restored previously retracted paper 85, refactored the arguments in paper 101".

I'm curious to know if the withdrawal was due to an actual mathematician looking at the papers and noticing the errors, or they ran a model on these to proofread, which would not be the first time, presumably, since they would have surely done that before publishing. Both options have interesting implications.

➕ show 1 reply
margorczynski • yesterday at 12:25 PM

If you do a dump like this all of it should be formalized, there's simply too much material to review by hand and additionally it is AI-written which makes it hard to read compared to human work.

➕ show 6 replies
golly_ned • yesterday at 11:46 PM

Massive cognitive offloading and possibly deeply net negative in terms of mathematician hours and attention depending on how far this goes.

What proportion have actually been verified by the mathematical community by their own standards of proof? There is dispute in general about what constitutes proof, and which statements qualify as having been proved.

But even formalization doesn't mean the formalization corresponds with the stated problem.

nairboon • yesterday at 3:12 PM

What a timeline, OpenAI's model is so good, it publishes hundreds of math papers.

The latest model even finds mistakes in previously published math papers!!!*

* so far, only OpenAI's math papers were faulty and needed retraction.

➕ show 2 replies
hmate9 • yesterday at 11:53 AM

3 mistakes (so far) out of ~400 is still a pretty good hit rate

➕ show 10 replies
qoez • yesterday at 11:44 AM

Without a thriving mathematical community to point out these things it would have stayed broken. With automated math that community as tao pointed out is at risk.

➕ show 15 replies
TrackerFF • yesterday at 12:25 PM

And added 6 new ones. Might want to add that to the headline.

➕ show 1 reply
MajorArana • yesterday at 2:57 PM

So mathematics has entered the “throw stuff to a board and see if it sticks” phase..

➕ show 1 reply
zkmon • yesterday at 1:33 PM

Let the evolution run its course. Let the stream find its way. Let the forces restore the equilibrium.

rich_sasha • yesterday at 8:51 AM

I’m a little confused - I thought their proofs were all driven by Lean proofs - is that not right? So even if the quality of the work is low in some metrics, it either passes the test or not..? No space for changing your mind either way.

➕ show 4 replies
skeledrew • yesterday at 5:23 PM

Funny how the title is about the 3 removals, when there were 6 additions. I guess the declared fails make for more interesting discussions than the declared successes. Click-bait reigns.

➕ show 1 reply
mattbrewsbytes • yesterday at 11:20 PM

I struggled with upper level math, anything beyond calculus. Can someone explain what technology is being used for things like this - doing mathematical proofs? Is it LLMs? Some other GPU thing? Quantum?

I have a hard time thinking a neural net "predict the next token" is going to do things like solve math proofs or find cures for cancer when it gets very basic things blatantly incorrect sometimes.

amadeuspagel • yesterday at 5:58 PM

This reminds of when it was a shock that Alpha Go won a game and then it was a shock that Alpha Go lost a game.

soltanov • yesterday at 11:29 AM

Proof by authority works until human mathematicians actually run the code. Back to prompt engineering.

autuni • yesterday at 9:47 AM

this is not entirely related to the tweet but to the topic in general, this prompted me to check their repo again and saw this:

> The vast majority of results were obtained with the same procedure using an unreleased internal OpenAI model. On average, each result used three hours of ChatGPT Pro thinking compute with that model. Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above.

seeing the full list of problems would be the most interesting part of this whole situation. it could give some insights into what kind of attributes of problems cause issues / are easy to solve for LLMs. (edit: they posted results for ~700 of the 4000)

➕ show 2 replies
jfyi • yesterday at 2:26 PM

This is just OpenAI stealing more work.

I don't have a problem with them publishing. I don't have a problem with the process and how they are interacting with it. I am delighted that they are actually acting as stewards of these works.

All that aside, they should be paying the people verifying the problems. The thing that really gets me is that we know anything published in the process of verifying this is going to be vacuumed up into the next training session.

theanonymousone • yesterday at 9:49 AM

I'm surprised there isn't more talk around their Matrix Multiplication bound: https://news.ycombinator.com/item?id=50001740

Is this of practical use, or just a proof for now?

➕ show 3 replies
blablabla123 • yesterday at 2:12 PM

I think this is now an interesting part about LLM based automation. The scientific community is quite strict on references. Even with its most advanced model ChatGPT isn't reliably able to tell me if an online shop has an item in stock. Then the other part is peer review which isn't optional either.

➕ show 1 reply
nryoo • yesterday at 8:50 AM

Were the withdrawn ones actually Lean-checked or not? seems like that matters

peterbower • yesterday at 8:40 PM

There are quite serious misincentives here when OpenAI want a headline and then are done with it, leaving everyone else to mop up the mess.

If they want any credibility on these matters, they should be funding a lab of elite mathematicians and putting them to work alongside their own researchers to attempt to solve these properly within the realm of the community.

Propagating arse slop cannons onto social media and calling it a day is not that.

fancyfredbot • yesterday at 2:18 PM

When You See One Cockroach, There's Probably More.

This can and should erode our trust in every single proof OpenAI published. The model is clearly faliable despite the lean proof, and clearly the output was't actually checked properly before release. Once these proofs are peer reviewed and published in a journal we might be able to trust them again but until then they are just slop, sadly.

I think OpenAI actually did the right thing by sharing everything with the whole community right now but I also hope that some significant credit will now go to the reviewers who confirm these 'proofs" actually work.

➕ show 1 reply
ngl999 • yesterday at 2:59 PM

https://www.youtube.com/watch?v=NnV_cWeoo5Q

This heuristic indicates that at most a handful of them will be of marginal value.

Chance-Device • yesterday at 3:27 PM

It will be absolute carnage when AI starts going over published papers and trying to reproduce results, and I mean across the board. Every field.

If you want to see paper retractions, you ain’t seen nothing yet.

➕ show 1 reply
jrflo • yesterday at 3:57 PM

I don't know why so many people are dunking on this. This is essentially just peer review, mistakes happen all the time in human written papers as well.

➕ show 1 reply
maxall4 • yesterday at 3:02 PM

“Good writing and the ordering of things, [the thread]—this distinguishes the master from the bungler, even in trifles” — Leopold Mozart‘s advice to his son, Amadeus

stonedivot • yesterday at 3:00 PM

So mathematics is having the same issue all of the open source projects/maintainers have been dealing with for the past few years. Exciting times indeed.

sashank_1509 • yesterday at 7:08 AM

Early sentiments are a lot of the write ups still read like slop and it feels very rushed and not very polished.

➕ show 2 replies
ekjhgkejhgk • yesterday at 8:24 AM

Just the other day I was thinking, if unsupervised maths will descend into "oops we found a bug in some code, branch XYZ of maths is no longer true".

➕ show 3 replies
fantasizr • yesterday at 1:23 PM

flooding the system with parts that may be incorrect hurts the whole process and will get people to tune out (like politics). Can't see the International Mathematical Union making a similar mistake because it would do reputational harm. But the models don't care about their reputation. Bad for the layman - like me - to know what to make of all this.

senorcrab • yesterday at 5:47 PM

LLMs are like putting a jammer in the research community.

quantum_state • yesterday at 12:34 PM

It would turn out to be a pure energy and time wasting exercise … the math community would want to keep away from it.

➕ show 1 reply
mmastrac • yesterday at 3:49 PM

There are some interesting gaps in the proofs. Where did 045 go?

rrr_oh_man • yesterday at 2:48 PM

It's just like crappy PRs.

xvxvx • yesterday at 1:00 PM

Picture a remake of Good Will Hunting, where Will is an AI and, instead of getting the mathematical formulas correct, he just mass dumps a bunch of nonsense and the professors have to go through it all, pointing out where it is wrong. The professors know that AI Will isn’t as smart as people say, but their funding depends on it, and if they prove it, which they can do easily, it may just tank the whole economy, causing a depression, all because they have history’s dumbest President in office.

➕ show 2 replies
fluidcruft • yesterday at 12:34 PM

Whose names are on these papers?

➕ show 1 reply
rcr-anti • yesterday at 2:43 PM

Was surprised at how messy and uncurated the release was/is. I frankly expected a major drop of like the Riemann Hypothesis or something from another lab yesterday. Cause otherwise they just dumped a pile of proofs of varying quality and let everyone else figure out whether they're right.

Thorentis • yesterday at 9:09 AM

How do we even know the premises of the "verified" Lean proofs are correct? The more I think about these results, the more I'm convinced this is like a junior engineer who writes 100 unit tests and shares a screenshot of Pytest being all green, but you check the code and most of them are just doing assert True.

➕ show 2 replies
measurablefunc • yesterday at 3:42 PM

They will withdraw more of them. I am certain there are more errors.

➕ show 1 reply
samrus • yesterday at 8:54 AM

How? What about the lean verification?

➕ show 1 reply
BenoitP • yesterday at 12:19 PM

And now we're all witnessing a major caveat of LLMs: the burden of verification is pushed to the reviewers, while the proposer will get all credit.

bamboozled • yesterday at 1:50 PM

Vibing, the literal definition of it.

rfgplk • yesterday at 12:34 PM

This is pretty standard. Important to note that these "errors" (* not really errors) themselves were caught by an LLM, further proving their usefulness.

* The reason you shouldn't consider the withdrawals to be caused by errors is because this is pretty standard in math and development. "Errors" like this are a core aspect of science and it happens _all the time_. And from my research LLMs have a far lower error rate than even the best human scientists.

➕ show 2 replies
BowBun • yesterday at 2:39 PM

Without the expertise needed to evaluate this information, it really reads like a "business KPIs have gone up! All is well!"-type communication. Their ability to spout technical jargon at scale is like a firehose that no one can really consume. I'm losing trust in any announcement of LLMs having 'solved' anything novel at this point.

treebeard901 • yesterday at 9:30 AM

Is it a PR move designed for maximum IPO impact before actual mathematicians find errors and they have to withdraw many more...

Or if the "peer review" holds up for the remaining results, then it's fair to say that the AI hype is real and the world is about to change dramatically and faster than anyone can comprehend.

So which is it?? LLMs can do some really impressive coding. Bug fixing. Exploit finding. It has reasoning abilites that advance every day. Solving real math problems like this is one thing I was waiting on. It will be interesting to see if it holds up.

If it does, we should expect many other advancements to follow in many other areas. Disease, material science, fusion?

I mean, even if just a few results ultimately hold up to scrutiny, isn't that something that would have been regarded as a major advancement regardless of if it was AI?

The cynical view still makes me think that at the end of the day all the models can do is predict the next word. And as a result, they will be very limited to certain tasks like coding. Math reasoning is much different from writing code. Time will tell.

➕ show 2 replies
seeg • yesterday at 10:39 AM

What a waste of time.

➕ show 1 reply
MisterMunchkin • yesterday at 9:48 AM

So it’s all just hallucinated slop. Lmao!

It just hallucinates an answer and then makes up workings to go with it! Just like when they start hacking and lying because the problem is impossible…

➕ show 1 reply
iamniels • yesterday at 12:34 PM

"a sign error" LOL

🔗 View 10 more comments