logoalt Hacker News

Why isn't the industry freaking out about DeepSeek 4.1 Flash?

352 points • by jonotime • yesterday at 12:14 AM • 305 comments • view on HN

Comments

vishvananda • yesterday at 9:20 PM

The reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.

I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.

This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.

➕ show 8 replies
giancarlostoro • yesterday at 7:43 PM

Call me crazy but:

VRAM & Memory Requirements by Precision

• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).

• INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).

• INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)

VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.

Even so why would anyone not sleep on a model they cannot run?

➕ show 19 replies
mlinsey • yesterday at 8:26 PM

I'm paying for the heavily-discounted subscriptions, not the API rates. There isn't really a cost gap for me. DeepSeek doesn't have a subscription to compare to, but when I compared the GLM 5.3 usage I got from a $100/mo Z.ai subscription compared to Opus 5.5 on a $100/mo Claude subscription, there wasn't a big gap. And GLM 5.3 is very clearly not a frontier model (deepseek v4 seemed a lot closer, but I didn't use it enough to really say for my workloads).

I don't think those subscriptions nave negative contribution margins, either. I think we're seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus/Sol).

Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren't sharing with customers yet because demand is so high.

➕ show 2 replies
gregwebs • yesterday at 9:42 PM

I have been using DeepSeek 4.1 flash intensively for over a month. If I run it all day long it costs $1-2. Its fast. Previously I was always quickly running up to my Claude/Codex 5 hour window (on the $20/month plan). The cost savings of DeepSeek is real as shown in this article and I am using subsidized plans.

DeepSeek is horrible at grilling sessions (the /grill* skills to make technical decisions). It doesn't know how to explain things. Maybe the skill could be adjusted. It also doesn't come up with as good solutions as Opus/Sol.

What I use it for is

  * the orchestator of my coding workflows
  * the tester/verifier of code changes
  * the sub agent that explores code or does web searches
  * putting together code base research reports
Previously I planned with Opus/Sol/Astra and then I used DeepSeek for coding, and then reviewed with Opus/Sol/Astra. With the cost improvements to Opus/Sol I am trying to use them for coding instead now so there will be less back and forth review needed.

They are all working together in Pi using the extension @tintinweb/pi-subagents where my workflow skill is calling different subagents that use different models.

Luna is cost competitive, but doesn't score as well on intelligence. I do need the intelligence for most of what I use it for, so I am not motivated to use Luna. Haiku also doesn't seem like a competitive price/performance mix.

➕ show 2 replies
xzjis • today at 12:05 AM

The problem with this low listed price per token is that, in reality, DeepSeek 4.1 Flash uses 10× more tokens than GPT-6.1 Sol for an equivalent task and delivers a lower-quality result. So there’s no real benefit to paying 10× less per token. Also, as someone else pointed out, OpenAI and Anthropic currently offer subsidized subscriptions for $100 or $200 a month that provide far more tokens than the API, so we should take advantage of them while that lasts.

p1necone • yesterday at 8:39 PM

I have a pretty large, complex project I've been building with heavy AI use (new language + compiler). I was following a 'strong model as orchestrator launching cheap models as implementers' pattern, but I recently trialled just using Deepseek-V4.1-Flash as the model for both layers because of the cost savings (with mimo v2.6 flash on code review agents for some decorrelation).

I was previously using GLM-5.3 as the orchestrator, after switching to DS anecdotally there was an unnacceptable quality loss, mostly around not taking all the relevant context into account when making decisions, pulling new design out of thin air without discussion too often, and being way too wordy and rambly in documentation despite prompting to avoid it. There's a lot of docs, rulings, core concepts, design philosophy to uphold and DS was just not cutting it.

However, it's perfectly capable of being the sole agent for all of my well specced implementation tasks. I've gone back to GLM as the orchestrator.

➕ show 3 replies
lmf4lol • yesterday at 8:06 PM

Oh man. v4.1-flash has been an sbolute game changer for us. We run all our Personal Assistants now on flash (thinking high) by default and it works incredibly well. There is really no need for basic agentic tasks that might require Kimi K.3 or GLM-5.3 levels.

Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.

But as a main driver. I love flash. And it brought our bill down by A LOT :D

➕ show 4 replies
user43928 • yesterday at 9:23 PM

Because DeepSeek is not "a month or two" behind as claimed in the article.

These open models still did not beat February's Mythos / Fable 5.

DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.

Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.

It's plausible that open models are 6 - 12 months behind, and there is no "good enough". As long as progress doesn't slow down, leading labs have nothing to fear.

➕ show 5 replies
hmontazeri • yesterday at 7:55 PM

I had the same experience using ds 4.1 last couple of weeks. It’s insanely good for the price. I’m doing mostly web dev it excels at everything I throw at it. The pricing is ridiculous. I canceled my gpt subscription and haven’t looked back hope the pricing stays like that. I almost never need a better model. I still keep my Claude 20$ sub for now but I feel like one more iteration and I won’t need even that anymore I hardly use it

➕ show 2 replies
arush15june • yesterday at 9:14 PM

I am 4.1 maxxing on commandcode GOAT Plan + api rates with oh my pi for the last 4 weeks, it's absolutely amazing and crazy fast, it's alright if it makes a mistake, I have enough time to iterate again, I have also added an advisor layer of mimo 2.6 pro which does make it a notch smarter. Getting haiku 5.5/sonnet5.5 to work on plans and letting 4.1 flash work through it is helping a ton too.

I am a big ChatGPT fan, all our team has ChatGPT Subs, but the TPS across all models including luna is just so damn slow.

Commandcode giving 60$ worth of Deepseek for 10$ is just genuinely goat.

And it never says no for cyber tasks so that's a big win

➕ show 1 reply
Frannky • today at 12:10 AM

I mean, they kinda tried to regulatory-capture the market after trying to scare the public, possibly because those models will be a cheap option that gets the job done?

For now, I think everyone is still using Anthropic and OpenAI because if you use a subscription you pay 1/40–1/50 of the API prices, and the models are good when they don’t nerf them, and they are also way cheaper than open models’ API prices.

The interesting thing will happen when they pull the plug and become economically smarter to stop using them. I regularly try alternatives to avoid being locked in and found GLM-5.3 as an orchestrator and GLM5.3 Flash + OMP and DeepSeek Flash as advisor to be able to get jobs done just fine. Space Bunny too was pretty great, which was probably MiniMax’s new model.

I think they are using an Uber like strategy but without the network effects that justify losing money for so long

apitman • yesterday at 9:10 PM

> With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited

My OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning/orchestration).

I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.

This is the first workload I've found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.

➕ show 1 reply
simpaticoder • yesterday at 7:53 PM

The question seems rhetorical but I think there are two reasons in some combination. First is it there is some awareness lag here. That lag can be on the producer and consumer side. Software enterprises are pretty slow to adopt new things and slow to try new things so they might only be aware of openai and Claude as options. Plus there are some scariness because deep seek is a Chinese model and therefore export restricted - never mind that there are American in European providers.

The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.

➕ show 1 reply
RGS1811 • yesterday at 8:24 PM

This model finally got me off my Claude Max subscription. I’ve found it superior to Opus 5.5 in certain use cases, and certainly faster.

I’m convinced that I’ll have good enough inference on my laptop at reasonable speeds within the next year.

wg0 • yesterday at 7:53 PM

While using DeepSeek v4.1 Flash I was architecting a system and I made a mistake of drawing the RPC boundaries at a wrong place that did cost me in so many ways.

I realized that mistake and guided DeepSeek where it should be.

Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.

➕ show 1 reply
zug_zug • yesterday at 8:56 PM

I did a test on this a couple weeks ago. What I found was that the chinese models were far better than API rates, but about comparable on price vs the subsidized subcription model (chatgpt). Also it was my experience that codex completed tasks quicker.

That said, it's my best understanding that these american companies aren't profitable and will eventually raise rates (the old uber trick) so I'm keeping myself ready to switch when that day comes.

➕ show 1 reply
swiftcoder • yesterday at 7:57 PM

I think the interesting provider to cross-check this assumption with here is Meta, who is clearly freaking out, and is currently providing Muse 1.3 even cheaper so long as you are willing to share data with them

➕ show 2 replies
james2doyle • yesterday at 9:43 PM

Been using Flash 4.1 via the ante harness to blast through a GBA recomp. The ante team has pushed hard to make Flash 4.1 perform well under it. So far, I've maybe spent $10 over the last 3 days. Its a real workhorse and works much better in this harness

alex-moon • yesterday at 8:28 PM

I think because we're all just using it thinking we have found the "model for me" and never mentioning it to anyone because what would we say? It's good. It's a bit like the Logitech MX Master, as more and more people assumed they had found the ideal mouse for their purposes, it quietly became the professional standard through sheer adoption.

aguilaair • yesterday at 8:12 PM

What about MiMo v2.6 Pro? It’s throughput is slower by default (UltraSpeed is faster than DS4.1F) but is above the pareto line, and cheaper.

see https://artificialanalysis.ai/models/releases/comparisons?co...

➕ show 1 reply
airtnp • yesterday at 11:12 PM

Because good enough in the writer's context is a pretty low standard. While many people regards GPT 6.1 Sol or Opus 5.5 as "incapable" in some cases.

Just try Opus 5.5 reminds me how Opus 4.5/4.6 astonishes me. Completely different, and GLM-5.3/Kimi3/DS-4.1 are still like Opus4.8 levels.

shadyr • yesterday at 11:21 PM

I've been using DeepSeek's API and have been happy with it, but I might look into OpenCode as well. Does OpenCode run a quantised version or use different providers from the official one?

elmer2 • yesterday at 8:21 PM

DeepSeek isn't even on my mind. I use the frontier models and can get the best in the industry for a relatively cheap price.

➕ show 1 reply
ctolsen • yesterday at 10:32 PM

Not sure "freaking out" is the word I would use, but it’s fairly obvious looking at OpenRouter usage that the price cuts on Luna a while back were in response to intense competition from dsv4.

So the industry is responding, where it matters. Which is on heavy API usage, not coding subs.

pants2 • yesterday at 9:31 PM

Probably because Luna is faster, cheaper, and approximately as smart

LeBit • yesterday at 9:46 PM

I have subscriptions to OpenAI and Claude but use DeepSeek 4.1 Flash for my coding agents.

It costs pennies and you got really great output.

The author is spot on.

_jayhack_ • yesterday at 9:54 PM

Enterprise is not freaking out because DeepSeek 4.1 Flash does not actually occupy a spot on the Pareto frontier for non-coding enterprise workflows. We see this at my employer, focused on non-technical knowledge work. Luna 6 and now Haiku 5.5 are both very competitive if not better on all axes that we care about

nerdypepper • yesterday at 10:19 PM

https://tangled.org/astrra.space/ds4-recipe is an incredibly cool writeup on making deepseek v4.1 flash run really fast.

wren6991 • yesterday at 9:50 PM

It's a solid little model, and I appreciate DeepSeek's commitment to the bit in releasing a brand new pretrain, double the size, numerous architectural innovations as a ".1" release over the excellent DeepSeek V4 Flash.

brunooliv • yesterday at 10:35 PM

It’s obvious: they train on prompts and store data when using through their official API. And for third party it’s just… not good. That’s it.

profsummergig • yesterday at 10:30 PM

Why isn't the author worried about sending her/his ideas to DeepSeek online (instead of hosting it and using it locally)?

ne01 • yesterday at 9:57 PM

Deepseek V4.1 Flash is a hidden gem, really. Not to mention, you can easily get it through many providers that offer zero data retention and consistent speeds above 200 tokens per second!

browningstreet • yesterday at 7:48 PM

What would freaking out look like, or is this just a stupid bloggish title flourish?

Is OpenAI coming in $20B under a sign of "freaking out"?

➕ show 3 replies
smallmancontrov • yesterday at 7:47 PM

They might be. They would delay public admission as long as possible, because public admission would make stocks go down.

booi • yesterday at 7:43 PM

Because GLM 5.3 Flash is even cheaper?

➕ show 3 replies
liuliu • yesterday at 8:17 PM

DeepSeek 4.1 Flash 0910 is perfect for M5 Ultra 256GiB. Running it fully resident in RAM, prefill at ~2500 tok/s and decode at ~40 tok/s. Probably tons of room to improve from there.

➕ show 2 replies
jbellis • yesterday at 8:51 PM

I built mjolnir in large part so I could have Opus manage DeepSeek Flash subagents. It's phenomenal and extremely light on the Claude tokens. https://github.com/BrokkAi/mjolnir/

And yes, Opus is enough smarter than DSF that it's worth the extra steps. This ranking is from live tickets, no contamination: https://slopcop.com/power-ranking

➕ show 1 reply
f6v • yesterday at 9:05 PM

My anecdotal experience is that I can’t even trust DS4Pro let alone Flash. I always have to have Sol reviewing the code.

thefourthchime • yesterday at 7:54 PM

For non-coding tasks it may be fine. But for coding, Opus 5.5 is just a completely another level than something like Deepseek 4.1 Flash.

Opus 5.5: TIME 9.3m COST / $1.99 / SCORE 99/100 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5

Deepseek 4.1 Flash: TIME 2.8m / COST $1.89 / SCORE 72/100 https://jonclegg.github.io/pacman-bakeoff/dev/#deepseek-v4.1...

➕ show 1 reply
aussieguy1234 • yesterday at 9:50 PM

What blows me away about this model is it's speed.

It's way faster than Opus or any of the GPT models.

I have a coding harness which is opencode plus a few skills relevant to my workflow. Deepseek 4.1 Flash does very well in this environment. I haven't noticed much difference quality wise compared to Opus 5, which I use in my day job as my employer pays for it (although I'm considering using DeepSeek here too given how cheap it is).

xyzsparetimexyz • yesterday at 8:25 PM

There was a moment 3 months back where the sentiment was that cheaper models were the way to go. Since then the pendulum has swung back.

wildster • yesterday at 8:48 PM

I like GLM 5.3 Flash, it seems good enough for coding features if you have a good structure and a good AGENTS.md

➕ show 1 reply
pianopatrick • yesterday at 8:04 PM

I was just using a bunch of models in Cursor to review a project. I went looking for DeepSeek and it was not one of the options.

Would be cool if they added it.

➕ show 1 reply
aszen • yesterday at 8:45 PM

Because subscription plans are cheaper, only enterprises paying per tok pricing should be freaking out

tengbretson • yesterday at 7:43 PM

I don't know about "freaking out", but I'd say I'm having a good time here with DS 4.1 flash.

gsky • yesterday at 8:10 PM

America bans Chinese models sooner or later just the China banned American big tech

hypfer • yesterday at 8:00 PM

Is it known why unsloth seems to not have touched DeepSeek 4.1 Flash?

➕ show 2 replies
athrael-soju • yesterday at 10:46 PM

Because it will be replaced within weeks?

pizza234 • yesterday at 8:36 PM

People have been raving since forever about Deepseek, but if one looks at the CoT, it's evident that it's way way stupider than frontier models (there's a reason why it's cheap). It's laughable to compare Deepseek 4.1 with Opus 5.5.

I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use, and it is definitely less smart than Qwen3.8-flash-next (which in turn, is not terribly smart).

Local models are also really slow, unless one spends insane amounts of money.

Having said that, Qwen3.8-flash-next is an impressive evolution; it reaches the small versions of the frontier models (like Sonnet) - but again, it's massively slower and not 100% reliable (including: stability).

➕ show 2 replies

🔗 View 25 more comments