I don't think anyone's disputing that the models are better than a year ago (though clearly all this, recursive self improvement stuff is utterly hypothetical and that's being generous, most of the progress has been on cost and fit and finish stuff). When it's on the plan, I'll use Fable for some stuff. If I'm paying for Opus? It's 4.5 or 4.6 which were dramatically better aligned and token efficient in the trace at a capability gap that's "you win some and lose some".
That can be true while it also being the case that to someone who has no idea how this stuff works under the hood, it's basically Dunning Kreuger in a box. The next person to go /u/PhdInEverything on me with Fable is getting an education in the history of hardware support for mixed precision training or something. Fable is a masterclass in refusal to ground and a dozen other alignment catastrophes.
So yeah, the models are still getting a little better, but from here out I think it's rapidly becoming a skill game.
[dead]