I don’t know how 5 can do so much better in benchmarks but absolutely suck to use in practice compared to 4.X. Fable feels better, Kimi and GLM also feel better sometimes but tbh all of them make plenty of annoying mistakes.
The prompts in most of the benchmarks match what you have to write to obtain good performance from Opus 5. Reading benchmarks is very revealing.
The prompts in most of the benchmarks match what you have to write to obtain good performance from Opus 5. Reading benchmarks is very revealing.