> There's no fundamental reason why a transformer with low bit weights would be inherently incapable of doing high dimensional function approximation,
Sure, in fact you can train a 1 bit model to approximate any function. But that is not the point. The problem is that the capacity per weight is diminishing below ~4.5bpw.
With QaT, can you compensate for that by adding more weights. So, instead of training a 1B x 4bpw model to capacity saturation you can train a 4B x 1bpw model and achieve the same capacity. But you will end up with the same number of bits in the model, just distributed differently.
Here are some older experiments of mine with very small models: https://github.com/cpldcpu/BitNetMCU/blob/main/docs/document...
I am quite certain that similar behavior is true also for larger models - it may be more difficult to experimentally demonstrate due to information bottlenecks and the flops required to train them to saturation.