logoalt Hacker News

seamossfetyesterday at 6:33 PM0 repliesview on HN

If you want to do a 1-bit model you have to QAT at pre-training with way more data than chinchilla to compensate for the cliffs (like 50x). Quantization on an existing pre-trained model will almost always collapse at 1-bit