I mean, this paper clearly shows that all models severely degrade for low bit quantization. Also, these experiments were performed on relatively old models that are not trained as close to capacity as newer ones.