logoalt Hacker News

Roark66today at 1:40 PM0 repliesview on HN

I'd rather buy two used rtx3090 than a single r9700 AI pro. More VRAM (some wasted due to it being non continuous), more RAM bandwidth, more aggregate compute.

Only if AMD made a card like this with 48G+ I'd consider it.

Also these 20-30t/s jumping to 150-200... Watch out for the massaged numbers coming from vendors.

I believe Intel has claimed something like 1400tok/s (generation! Not prefill) of Qwen3.6-moe on Arc b70.

I was actually very interested in this so I checked the details. Turns out it was 200 simultaneous users running the same 1024 token prompt :D so all the experts got maximum parallelism.

How often are you going to run 200 parallel sessions with a tiny context and same prompt running at 7tok/s.

Based on how much my rtx3090 is getting on a single user (150tok/s) I'm estimating b70 to probably get less than that.

Sadly nvidia is king now.

Also, most of us already have nvidia cards and no inference software supports mixing let's say nvidia, Intel and amd cards in inference of one model.