logoalt Hacker News

verdvermyesterday at 9:18 PM1 replyview on HN

This is the best side-by-side comparison I've seen so far, we're still waiting for it to become available for download, then we should get much better analyses.

https://news.ycombinator.com/item?id=48999291

I do agree that the benchmarks do not tell much of the story, but this applies to the closed models as well ime.


Replies

braeboyesterday at 9:35 PM

I’m not kidding when I say bullshit bench and simple bench are the only benchmarks that reflect real world utility in a way that matches my hundreds of hours of experience with frontier models: https://petergpt.github.io/bullshit-benchmark/viewer/index.v...

That said, while I find it hard to trust fireworks given their conflict of interest, their article is pretty good.

show 1 reply