logoalt Hacker News

DeepSeek V4 Flash on a Single AMD MI300X

338 pointsby zhoutongtoday at 10:00 AM86 commentsview on HN

Comments

majketoday at 10:44 AM

I don't think you can buy a single "MI300X" unit, right? Only the box with x8 of these at a cost of ~250K EUR.

show 4 replies
fergusfinntoday at 2:27 PM

nice! i think the higher HBM on Mi300x is really useful for this kind of thing

we did some work on this for 2xMi300x (kindly referenced in the readme) https://blog.doubleword.ai/deepseek-v4-flash-mi300x. https://hotaisle.xyz/quick-start hotaisle is great for getting Mi300x to experiment with

GTPtoday at 1:42 PM

Strange that in the prior art they didn't list DwarfStar, as it is able to run the same model (probably quantized differently though) in less memory. Maybe the author isn't aware of it?

show 1 reply
Tepixtoday at 1:28 PM

Unfortunately, the MI300X is an OAM module. The MI350P is the one you want: It's a PCIe card, but it has less memory: 144GB.

Luckily, DeepSeek V4 Flash will run in 144GB too because it's 256 MoE exports are native MXFP4 quantized.

show 2 replies
WhitneyLandtoday at 2:16 PM

Another headline of “model runs on x”, which usually means “let’s list how much you give up to run on x”.

Dumbed down quantization?

No. Full intended inference weights preserved, so far so good.

Slow performance?

No again. Looks like you could get over 150 tokens/second.

Give up context window size?

Yes. Original model is trained for and served at 1M, this is 256k. A very practical tradeoff though. Codex is in this range, and quality does start to drop off toward the full size.

show 2 replies
xorfishtoday at 12:51 PM

This is still quite a bit away from the performance that deepseek gets on their H800. In their DSpark paper they report a throughput of 15k tokens/s/gpu. The MI300 should be able to compete with the H800 so there are probably still quite a few optimizations that can be made.

show 1 reply
sylwaretoday at 1:41 PM

Is their hardware programming interface reasonable for implementing inference of frontier models: no quantization, several tera params?

BTW, how many many params open weight frontier models have? A few teras, 100s of teras?

show 2 replies
pop3zxcvtoday at 5:36 PM

[dead]

jkwangtoday at 11:03 AM

[flagged]

hn0tdqaek4today at 12:52 PM

[dead]

PrimeAlitoday at 5:55 PM

Great