logoalt Hacker News

techjamietoday at 4:44 PM4 repliesview on HN

With the performance gains they're claiming, I wonder if they implemented the Casual Encoder-Decoder technology from DeepSeek 4.1's paper.

I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.

How it works: https://miraflow.ai/blog/deepseek-v4-1-flash-causal-encoder-...


Replies

stri8tedtoday at 5:39 PM

This model was likely trained months before deepseek released their paper.

show 1 reply
Balinarestoday at 7:24 PM

Unless they already have something similar of their own, which is always possible, they'd be stupid not to. I don't suppose we'll ever know, though. It would not be a good look if after the trillions of dollars that have been thrown at US labs, investors found out that they're down to copying Chinese tech.

ACCount39today at 5:48 PM

I don't think it's particularly relevant?

They might be using something like this, or they might be using some other "increased sparsity" techniques, of which there are a great many. They also might be optimizing for something else - like less RAM use for KV cache.

Alternatively, they might be cutting into their margins and dropping the price because of stiffer competition from Astra. I do think that's unlikely though.

ryanggtoday at 5:10 PM

Getting a 403 on that link. Mind checking it once?

show 3 replies