logoalt Hacker News

eliyesterday at 9:49 PM1 replyview on HN

That's a plausible explanation but I'm not seeing evidence for it.

I have a personal benchmark suite of 14 real, non-public tasks. Opus 5 and Fable tied on 10, Opus won on 3, and Fable won on 1. It's a really strong model.


Replies

matt2000yesterday at 10:09 PM

This is an interesting idea, without giving away your benchmarks specifically what kind of stuff do you test? I might try to assemble something myself, it's so hard to determine model quality from the system card these days.

show 1 reply