logoalt Hacker News

tintor • yesterday at 7:50 PM • 0 replies • view on HN

Compression doesn't really work for model weights.

Model quantization and model distillation are two techniques to reduce model size.