logoalt Hacker News

Tepix • today at 1:55 PM • 4 replies • view on HN

Q2 quantization. Not interested.


Replies

tcdent • today at 2:19 PM

All of these projects targeting low spec systems and "100 tok/s" are the same 2 bit quant without much else. Conveniently none of them include any mention of accuracy in their published numbers. 4 bit is the floor.

➕ show 2 replies
ivanjermakov • today at 3:34 PM

These "revelations" are getting closer and closer to "download RAM for free" each day.

latentsea • today at 4:19 PM

You can run IQ3_XXS, IQ3_S, and IQ4_XS too. I've switched to IQ3_XXS and am running at 60 t/s on Strata vs the 21 t/s I was getting in llama.cpp. Better outputs too.

CamperBob2 • today at 4:54 PM

Larger models can tolerate Q2 quantization surprisingly well, especially if they were trained with quantization in mind. I don't know about 3.8 125B, but for example, there are 2-bit quants of Kimi K3 that exhibit strong reasoning and maintain decent coherence at longer contexts.