Out of curiosity, do you have any theories of why it works so well at such aggressive quantization l...

someone13 • today at 2:31 AM • 1 reply • view on HN

Out of curiosity, do you have any theories of why it works so well at such aggressive quantization levels?

Replies

It's a mix of extreme sparsity but with the routed expert doing a non trivial amount of work (and it is q8), and projections and routing not being quantized as well. Also the fact it's a QAT model must have a role I guess, and I quantized routed experts out layers with Q2 instead of IQ2_XXS to retain quality.

alt Hacker News

Replies