To add: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RC... IQ3_XXS is within a point of the fully unquantized model and IQ3_S actually beats the unquantized model on many tasks! You lose absolutely nothing. It is quantization magic :)
No magic needed, the quantization found the expert which deals in java and deleted it, so the model overall became better.
No magic needed, the quantization found the expert which deals in java and deleted it, so the model overall became better.