The tool I've been using, llm-compressor, can quant models that do not fit in memory (use the sequential pipeline)
https://github.com/vllm-project/llm-compressor
my setup to help you on your way: https://github.com/verdverm/quantr
Though it seems these will not be needed as Poolside has published quants & dflash with their models.