I (author of the original post and paper) can add a few things here: 1. My previous approaches with GPT 5.5 were really not very sophisticated in terms of my input. I threw the problem at it, and just kept encouraging it to go iterate through ideas without any success. 2. The approaches that are in the prompt, though they will seem cryptic to someone not in the field, are relatively natural ideas. In fact, the construction that worked was something that even 5.5 initially looked at but was just too weak to see how to make it work. From my view, I would have never gotten this result myself. Imagine you are telling a contractor to build the empire state building, and you say: "You should explore approaches that can include building materials like steel, wood, concrete, or clay, and any combinations of those. You can use arcs, columns, supportive beams, and anything else you can think of to solve load-bearing issues. Do not stop until you've completed a viable plan to construct the empire state building." And then the contractor shows you the finished empire state building using reinforced concrete and steel beams with all kinds of crazy ways of making everything stable; that's kinda how I feel.
There is a firm opinion in some circles that an AI can never do anything but recombine already known things despite a rising number of cases where that very much appears to be an inaccurate view.
Any slim possibility that vital information was given to the AI to make the task just one of recombination, rather than coming up with anything original, is grabbed with both hands no matter how tenuous.
I can understand why, nobody likes the idea of a machine capable of doing the job of a human, and jobs requiring the mind are especially sacrosanct. But like farm workers in the early part of the C20th, not liking the incoming technology won't change the effect it will have when companies get the idea they can make more money and hire less people.
Thank you for replying and for reporting your results. My schtick is AI (i.e. I publish in AI conferences and journals), not maths and I'm interested in the question of autonomy from a professional point of view: how close are we to a machine that can carry out the job of a mathematician like yourself, by itself?
If you, e.g. got an LLM (any one) to one-shot the problem with a prompt that said "solve this problem", then we're much closer to that, than if you spent a year trying with different LLMs and then finally got it to work with a lot of hand-holding and even suggesting the ultimate solution. An autonomous system can't succeed once a year, if it's going to be of any use. It should also be able to identify the right tools on its own, not rely on a human to tell it what to do.
This should also go to address some of the questions you pose on reddit, about the future of mathematics. If LLMs can already do your job fully autonomously (as I would explain the term) then ... you're out of a job. You and all other mathematicians, young or old.
I personally don't think we're there yet.
I hear what you say, btw, about never being able to get there by yourself etc. Maybe you would, maybe you wouldn't. What we know is you used a tool to get there, in fact a series of tools, and it took you many tries before you did. The fact that an earlier LLM tried and failed is interesting, but that may just mean you were capable of crafting a better prompt after your interaction with the earlier LLMs.