logoalt Hacker News

bitexploder • yesterday at 4:27 PM • 1 reply • view on HN

To add: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RC... IQ3_XXS is within a point of the fully unquantized model and IQ3_S actually beats the unquantized model on many tasks! You lose absolutely nothing. It is quantization magic :)


Replies

Muromec • today at 12:05 AM

No magic needed, the quantization found the expert which deals in java and deleted it, so the model overall became better.

➕ show 1 reply