All the benchmarks in the world don't matter if the model just straight up refuses to do mundane things. Claude has too much of an attitude.
I'm a kernel engineer. Fable 5 refused all my requests, falling back to Opus 4.8. My wife is a chemist. Her experience wasn't much better.
The only company to use Claude.md instead of Agents.md standard
I notably had an issue that it wouldn't work on a "remote execution" (running a command over SSH) coding problem until I did a sed to remove the word "execution". Incredibly dumb. I'm not doing any murders. Easiest to just switch to the Chinese models.
I think a lot of CTOs that signed enterprise contracts with Anthropic are going to be in for a rude surprise.
It's one thing to generate some code and ship it, but it's another when your developers don't understand said code and it brings down production. If the model refuses to assist debugging the problem because it triggers some safety mechanism, you might be fucked.
All the benchmarks in the world don’t matter if the subscription forces you into a walled garden of slopcoded apps. I’ll stick with Codex and, increasingly, open source SOTA models.