If I were to anthromorphize my experience with frontier models, it would be as a mentally challenged child with complete memorization of an encyclopedia and thesaurus.
Memorizing an encyclopedia is not going to help you solve open high-level math problems, is it?
> It has the ability to rapidly experiment and potentially succeed at tasks through trial-and-error, but not without constantly corralling it in the correct direction
It kind of could, given the above. Part of an LLM's advantage is that, much like a calculator or a Chess engine, it can iterate over a finite problem space far, far faster than a human can. That much is expected of a useful computing tool.
It is worth noting that we know literally nothing about how the counterexample was achieved. Technically speaking, the person who tweeted it could have solved it themselves with zero LLM assistance and then attributed it to Fable to boost their IPO and ensuing payday. I'm not saying that's actually what happened, but it's hard to draw conclusions without any transparency about the degree of human involvement.
My daily experience certainly does not reflect that of prompting a superhuman intelligence when it routinely flubs commands and destructively drops my PATH, or bypasses an instruction about passing tests by burning millions of tokens constructing a completely new test suite that rubberstamps its own work when it can't pass the real tests.