I think treatment is not good benchmark - it requires lots of waiting and lots of regulatory work. The better benchmark - in my opinion- would be math discovery.
That is an absolutely terrible benchmark. Inference over a bounded search space is not a good measure of what "intelligence" actually is. part of the reason they are using math and not something actually challenging like long distance interstate trucking is because it's so much simpler and easier than what make intelligence intelligent.
Claude Fable recently proved the existence of complex structures over S^6 (6-sphere).
If I had to guess, I think LLMs will be inventing highly original new mathematics within the next year. I think it will be approached as an optimisation problem, targeting how quickly LLMs can solve classes of maths problems as a function of the definitions they need to conjure up to do so.
My point was that we won't care about benchmarks anymore because we would see an obvious and completely unprecedent increase in productivity (and I believe it will likely come from the same people who will develope such machine).
The reason most of the conversations are focused on benchmarks is because we are still in the age of weak AI.