if you have the VRAM, use offical release. quantized model lose focus after long context and can do damages or thinking loop