Those 27.3% are still in the ballpark of modern models:
- Sonnet 5 - 12.4%
- Luna - 17.3%
- Grok 4.6 - 20.3%
- Sol - 37.3%
- GLM 5.3 - 41.8%
- Opus 5 - 51.8%
Astra is 58%. The current title says it's "rivaling Astra"
DeepSeek v4.1 Flash 31.2%
Source: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash#compa...
This explains a lot about Sonnet 5.
So, better than Sonnet and Luna? lol
GLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or...
Also a lot of questions to benchmark because opus 5 is completely useless model right now.
I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?