logoalt Hacker News

stingraycharlesyesterday at 10:57 PM2 repliesview on HN

As always, benchmarks rarely paint the whole picture. It also seems like this article is somewhat biased, eg when Fable and Kimi are close but Fable wins it’s “dead heat”, but when Kimi wins it’s “Kimi wins”. GPT 5.6 seems to be missing as well.

I am really eager to give Kimi K3 a try, but I’ll reserve my judgement until I’ve worked with it for at least a few days.


Replies

rogerrogerryesterday at 11:13 PM

The apparent bias may be explainable as it’s not remarkable for OpenAI or Anthropic to be slightly ahead. It _is_ remarkable for an open weights model to be better than the closed models from the trillion dollar (allegedly) companies.

show 1 reply
jmtullosstoday at 6:27 AM

it appears that 3% is the threshold as 3.1% gets the nod the other way