I don't disagree, but that isn't really a quant anymore. That's just training a new model imo.
My understanding is that most of these quants - to achieve better precision than just naively dropping bits - require some finetuning against the activations of the higher precision model, so the line already seems kind of blurry.
My understanding is that most of these quants - to achieve better precision than just naively dropping bits - require some finetuning against the activations of the higher precision model, so the line already seems kind of blurry.