Quantization is typically very cheap and fast. It can even be done on hardware that does not fit the model, by processing the weights layer by layer.
I use this project: https://github.com/vllm-project/llm-compressor