logoalt Hacker News

user43928 • yesterday at 9:23 PM • 5 replies • view on HN

Because DeepSeek is not "a month or two" behind as claimed in the article.

These open models still did not beat February's Mythos / Fable 5.

DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.

Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.

It's plausible that open models are 6 - 12 months behind, and there is no "good enough". As long as progress doesn't slow down, leading labs have nothing to fear.


Replies

BobbyJo • yesterday at 9:29 PM

I was thinking about this earlier today and I came to the following question:

If you had a model 10x as capable as the best model out today, but it cost 100x more, would there be a market, and, if so, how big?

I think there would be a market and I think it would be large.

So, I agree.

➕ show 2 replies
runtime_terror • yesterday at 11:32 PM

Perhaps on certain benchmarks and for certain work, but anecdotally I've not been able to see a difference between it and Opus on a lot of dev work (web, Go, iOS/AppleTV native, scripting, general tasks)

airtnp • yesterday at 11:13 PM

Agree, will see after the dilution and CoT hack fixed, will they keep the pace now. MiMo had some good numbers recently because it's discovered that the post evaluation RL directly exposes answers to models, so RL and evaluation is runied.

aleqs • yesterday at 10:03 PM

that just sounds like openai/anthropic cope/propaganda, based on absolutely nothing objective lol

even their harnesses are far surpassed by pi and opencode at this point

also sick 'rumors' lmao, apparently marketing through rumors is in vogue these days

➕ show 1 reply
ForHackernews • yesterday at 9:44 PM

There absolutely is "good enough" and I agree with this author: DeepSeek 4.1 Flash is plenty good enough for all the things I would trust an AI to do at my job.

➕ show 1 reply