Easy enough to explain: they're benchmaxxing. Fable is intelligent but not benchmaxxed. Opus is less intelligent but benchmaxxed.
That's a plausible explanation but I'm not seeing evidence for it.
I have a personal benchmark suite of 14 real, non-public tasks. Opus 5 and Fable tied on 10, Opus won on 3, and Fable won on 1. It's a really strong model.
Yep , same with 5.6. Fable is still the best.
Honestly that's the simplest explanation and thus likely the correct one.
That's a plausible explanation but I'm not seeing evidence for it.
I have a personal benchmark suite of 14 real, non-public tasks. Opus 5 and Fable tied on 10, Opus won on 3, and Fable won on 1. It's a really strong model.