logoalt Hacker News

egerestoday at 7:27 AM4 repliesview on HN

It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 (https://artificialanalysis.ai/models/deepseek-v4-1-flash) gets 39. According to the appendix at the bottom of https://mimo.xiaomi.com/mimo-v2-6 the deepseek model sometimes surpasses mimo and it's not so far behind in capabilities. A week ago opus 5 appeared 1 points ahead of fable 5 despite fable being a much smarter model (this has been corrected already)


Replies

SyneRydertoday at 9:17 AM

The main AA benchmark keeps changing, and had to be radically changed when Astra came out and showed zero improvement over GPT 5.6 Sol in their benchmark. Opus 5 is still 1 point ahead of Fable 5.0 on the index, if you manually add Fable 5.0 back into the list, so it hasn't actually been "corrected". It's only Fable 5.1 that is shown as ahead of Opus 5.

The AA benchmark is a weighted average of other benchmarks and some internal ones. I think the difficult part is finding benchmarks that reflect your own use of the models.

show 2 replies
GodelNumberingtoday at 9:05 AM

> It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 gets 39.

Why?

show 1 reply
conceptiontoday at 12:59 PM

DS 4.1 is good but it’s clearly not as “smart” as non-flash models- it just doesn’t have the training data. Without a solid plan, it goes off the rails pretty regularly.

Shekelphiletoday at 3:41 PM

All of the chinese labs have been overfitting on benchmark data to game the results for a while now - MiMo and Deepseek are not anywhere near frontier and mostly compete with models like Luna - which they are still worse than.

There isn't much compelling reason to use these unless you are just averse to giving money to openai/altman. A $20 codex sub gives you ~$150 of luna use per weekly limit, while there isn't any good subsidized options for chinese models at all (and the few who were subsidizing, like opencode, rugpulled by reducing monthly limit to $60 to $15 with no notice to users).

show 2 replies