Counter evidence:
https://arxiv.org/abs/2603.00042
It is true that block-headed quantization of everything doesn't work. But, as I am sure will read, if you remove the spiky parts of the parameter values, the residue can be dramatically quantized and compressed while retaining performance.
This is a multi-modal sort of compression where you use different techniques for different phenomena. Simply compression all of the weights, each in isolation, ignores the benefits of compressing them collectively.
I mean, this paper clearly shows that all models severely degrade for low bit quantization. Also, these experiments were performed on relatively old models that are not trained as close to capacity as newer ones.
The capacity of your model scales with total number of bits in your weights. Shannon entropy.
Yes, there are many way to distribute the error. And LLM are usually not trained to capacity or are prevented from this due to information bottlenecks in the gradiant flow. This allows unlocking wasted capacity by distribtion of the quantization error. But I am pretty certain that there is no way to trick Shannon.