3 mistakes (so far) out of ~400 is still a pretty good hit rate
I would say not. For a mathematicians, having to retract more than 2 papers in a lifetime is already a big issue in their career.
Extrapolating the rate, there won't be any left by the end of next quarter. I jest, but reading these is arduous and finding holes in them is going to take time for anyone daring to.
I bet it's a lower rate of errors than the typical human math paper of 2025
Exactly, I was glad to see these withdrawals, its a natural part of a healthy ecosystem of scientific review, hypothesis, claim, test, refute, extend, withdraw, its the heart of science.
IMO If you take out all the stupid human aspects mostly related to fear, egos, etc, we should brace the imperfect and helpful tools, whatever they are, improve them so they are as easy as possible to review, and keep that core scientific discovery loop going
well, it might be that these proofs are correct or it might be that people aren't bothering to spend a lot of time checking whether they are correct. OpenAI already has a pretty bad reputation in the mathematics community for how they are approaching this process, they seem to be more interested in creating a story for their IPO than advancing math.
But can the others even be "disproven", given that they apparently are so messy and awful that no humans can follow them? Shouldn't the onus instead be on OpenAI to prove that they're right, instead of hundreds of mathematicians wading through slop?
It has literally not even been a day yet
How many mathematicians need to retract ~1% of their papers?
Except Mathematics and science is not done this way, its not code that you can just release bugfixes to, its not a numbers game, its about furthering our shared knowladge. If it is done by flooding everything with a bunch of paper that have not been peer reviewed and verified, and are known to be error prone, that just takes away a bunch of mental capacity from scientists, and time, to manually verify all 400 of them. The issue here is how openAI approaches the science, not their hitrate.
I’m conflicted. I guess we’ll see what the final total is once an enormous level of unpaid human effort is expended verifying the AI outputs. A little sad if that’s the future of math.
It kind of reminds me of when tech giants open source a project as a means of putting a positive spin on abandonware. “Here’s the source! Any problems are yours to fix now. You’re welcome”