logoalt Hacker News

skolosyesterday at 7:22 PM0 repliesview on HN

There was a study specifically related to Qwen3.8 27B that showed that kv cache quantization has almost no impact on this model all the way to q4:

https://arxiv.org/html/2609.04098