logoalt Hacker News

nbutton762 • yesterday at 7:48 PM • 2 replies • view on HN

First I just want to say, specifically for Post Training Quantization, I absolutely agree with you. As for why 4 bit is the limit, the ParetoQ paper has some interesting theories: https://arxiv.org/abs/2502.02631

That being said, there's a slight misconception about models at lower than 4bpw. There's no fundamental reason why a transformer with low bit weights would be inherently incapable of doing high dimensional function approximation, but training a model at one precision and then quantizing to a lower precision means the training loss is never calculated based on the quantized state.

There's a huge difference between "I trained a ternary model from scratch to do X" and "I trained a model at fp16 to do X and then squashed the hell out of it". Quantization Aware Training is the solve, but it's really expensive compared to a one-time, offline translation of existing weights.


Replies

cpldcpu • yesterday at 9:39 PM

> There's no fundamental reason why a transformer with low bit weights would be inherently incapable of doing high dimensional function approximation,

Sure, in fact you can train a 1 bit model to approximate any function. But that is not the point. The problem is that the capacity per weight is diminishing below ~4.5bpw.

With QaT, can you compensate for that by adding more weights. So, instead of training a 1B x 4bpw model to capacity saturation you can train a 4B x 1bpw model and achieve the same capacity. But you will end up with the same number of bits in the model, just distributed differently.

Here are some older experiments of mine with very small models: https://github.com/cpldcpu/BitNetMCU/blob/main/docs/document...

I am quite certain that similar behavior is true also for larger models - it may be more difficult to experimentally demonstrate due to information bottlenecks and the flops required to train them to saturation.

cpldcpu • yesterday at 10:02 PM

>As for why 4 bit is the limit, the ParetoQ paper has some interesting theories: https://arxiv.org/abs/2502.02631

Its not so clear to me what their real explanation is. Yes, it is possible to redistribute the quantization error, but this only works when the model is not trained to capacity.