logoalt Hacker News

fallingbanannayesterday at 5:15 PM5 repliesview on HN

Those 27.3% are still in the ballpark of modern models:

- Sonnet 5 - 12.4%

- Luna - 17.3%

- Grok 4.6 - 20.3%

- Sol - 37.3%

- GLM 5.3 - 41.8%

- Opus 5 - 51.8%


Replies

mokreyesterday at 9:04 PM

GLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or...

Also a lot of questions to benchmark because opus 5 is completely useless model right now.

I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?

p1eskyesterday at 7:46 PM

Astra is 58%. The current title says it's "rivaling Astra"

show 1 reply
Readeriumyesterday at 6:30 PM

DeepSeek v4.1 Flash 31.2%

Source: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash#compa...

show 1 reply
nijaveyesterday at 8:32 PM

This explains a lot about Sonnet 5.

airstrikeyesterday at 9:24 PM

So, better than Sonnet and Luna? lol