logoalt Hacker News

kittikittitoday at 6:52 PM0 repliesview on HN

I quantize my models with llama.cpp and it's usually one command. Some of their quants are fine-tuned by architecture but it's only to squeeze out every little performance benefit.