Have you looked at using oMLX?
https://omlx.ai/
Would second this, I switched to oMLX I get ~75 tok/s on Qwen3.6-35B-A3B-4bit on a 48GB M5 Pro
Would second this, I switched to oMLX I get ~75 tok/s on Qwen3.6-35B-A3B-4bit on a 48GB M5 Pro