The capacity of your model scales with total number of bits in your weights. Shannon entropy.
Yes, there are many way to distribute the error. And LLM are usually not trained to capacity or are prevented from this due to information bottlenecks in the gradiant flow. This allows unlocking wasted capacity by distribtion of the quantization error. But I am pretty certain that there is no way to trick Shannon.