> Initially, the rumors about these OpenAI results had it that there were solutions to 400 problems, and now the rumors are that further dumps are arriving soon.
That is mental if true. Though if once again most of them will lack a Lean proof, bit hard to feel confident about their correctness. They might trust their model, but given the community impressions on manuscript prose quality, I have a hard time imagining they are any better than its coding output, where closed loop verification continues to be essential. OpenAI might have great trust in their internal model, but this seems like a needlessly hazardous way test that trust out. The optics of those retracted papers from a few days ago were bad enough already.
I have separately heard of efforts regarding Lean proof optimization. Since it's all self machine checkable, might be valuable to do before a release like this too. I'd imagine a shorter Lean proof means a shorter manuscript too.
I read somewhere that 56% of the problems had some Lean Proof, but upon inspection, there are gaps between what Lean proves and what English says about the proof.
Making an English language description of a problem and proof steps match the Lean proof is apparently harder than the halting problem and is not something that a Lean proof alone or the current dump or even the Navier-Stokes dump solves.
A good breakdown of the problem is here: https://arxiv.org/html/2610.08144v1 or also see here: https://terrytao.wordpress.com/2026/10/09/what-mathematician...