>As for why 4 bit is the limit, the ParetoQ paper has some interesting theories: https://arxiv.org/abs/2502.02631
Its not so clear to me what their real explanation is. Yes, it is possible to redistribute the quantization error, but this only works when the model is not trained to capacity.