logoalt Hacker News

square_usualyesterday at 5:56 PM3 repliesview on HN

Easy enough to explain: they're benchmaxxing. Fable is intelligent but not benchmaxxed. Opus is less intelligent but benchmaxxed.


Replies

eliyesterday at 9:49 PM

That's a plausible explanation but I'm not seeing evidence for it.

I have a personal benchmark suite of 14 real, non-public tasks. Opus 5 and Fable tied on 10, Opus won on 3, and Fable won on 1. It's a really strong model.

show 1 reply
llelouchyesterday at 6:03 PM

Yep , same with 5.6. Fable is still the best.

show 1 reply
6thbityesterday at 7:44 PM

Honestly that's the simplest explanation and thus likely the correct one.

show 1 reply