logoalt Hacker News

cpldcpu • yesterday at 6:41 PM • 4 replies • view on HN

I understand the obsession with low bit quantization, but it is empirically quite evident that it is not possible to compress models to less than 4 bit per weight without severe loss of capabilities².

It may be nice as an experiment, but it is obviously a very inefficient route for model training: spending all the flops on a saturated model only to prune its capabilties.

²As to why, I have seen few explanations. But the empirical evidence is there.


Replies

nbutton762 • yesterday at 7:48 PM

First I just want to say, specifically for Post Training Quantization, I absolutely agree with you. As for why 4 bit is the limit, the ParetoQ paper has some interesting theories: https://arxiv.org/abs/2502.02631

That being said, there's a slight misconception about models at lower than 4bpw. There's no fundamental reason why a transformer with low bit weights would be inherently incapable of doing high dimensional function approximation, but training a model at one precision and then quantizing to a lower precision means the training loss is never calculated based on the quantized state.

There's a huge difference between "I trained a ternary model from scratch to do X" and "I trained a model at fp16 to do X and then squashed the hell out of it". Quantization Aware Training is the solve, but it's really expensive compared to a one-time, offline translation of existing weights.

tempoponet • yesterday at 7:01 PM

Influencers have latched onto the pitch that everyday people can run frontier models on an 8gb GPU while sticking it to the labs. There's a large audience of people who haven't had the hardware to test larger quants to see the difference.

They would be better served with smaller models that can reliably call tools and generate structured outputs without looping or totally hallucinating. These projects exist, but aren't getting amplified.

➕ show 1 reply
ted_dunning • yesterday at 7:41 PM

Counter evidence:

https://arxiv.org/abs/2603.00042

It is true that block-headed quantization of everything doesn't work. But, as I am sure will read, if you remove the spiky parts of the parameter values, the residue can be dramatically quantized and compressed while retaining performance.

This is a multi-modal sort of compression where you use different techniques for different phenomena. Simply compression all of the weights, each in isolation, ignores the benefits of compressing them collectively.

roosterIllusi0n • yesterday at 8:41 PM

There are mixed weight models that make it work. I have been using qwen3.8_27B_UD_Q3_K_XL. This does 46 tok/s on a 5080 with 16gb vram and it can pretty much do any systems work because it can test to make sure it did it right.