My understanding is that most of these quants - to achieve better precision than just naively dropping bits - require some finetuning against the activations of the higher precision model, so the line already seems kind of blurry.