Q2 quantization is basically giving a capable model a lobotomy. It will not accurately represent how smart or capable something like qwen 3.8 27B in Q8 will be.
Sure, but this is true for all lossy compression (audio, images, etc.)
Given 16GB of VRAM, what will give me the best experience in OpenCode? Currently using Qwen3.8_Q_3
Sure, but this is true for all lossy compression (audio, images, etc.)
Given 16GB of VRAM, what will give me the best experience in OpenCode? Currently using Qwen3.8_Q_3