logoalt Hacker News

ai_ja_nai • today at 2:37 PM • 2 replies • view on HN

I am not getting it: I see a fp2 quantized model going on a 5090 with 64GB of RAM at 90 tops with -10% accuracy over original model. How is this supportive of the claims?


Replies

kennywinker • today at 6:20 PM

Yeah, -10% accuracy (probably more like -20% in reality) sucks, but only if you could be running it at 100%.

That's the exciting part of this - before the best you could run on <24gb vram was qwen3.8-27b at q4 quantization. Now you can run a nerfed 125B parameter model on under $800 of hardware, and it beats a less-nerfed 27b model.

ai_ja_nai • today at 2:39 PM

(64GB not VRAM, I meant) I also see people claiming fast performance on a 128GB machine, which is not exactly consumer hardware)