Ok so i took the time to benchmark this on the following on my m5 max 128gb:
Qwen3-4B-Instruct-2507-4bit Qwen3.5-35B-A3B-4bit Qwen3.5-9B-MLX-4bit Qwen3-Reranker-0.6B-4bit Qwen3-Coder-30B-A3B-Instruct-4bit qwen2.5:0.5b
and the results are what i kinda expected to begin with, this adds next to nothing? - also the repo was pivoted from a playwright sub assembly to this not long ago - so my conclusion - THIS MIGHT be worth some watching if you have a model where no-one!, has optimized it at all - and where it does not use anything native to your platform.
results (averages)
VIA rapid-mlx :8902 (MLX) Decode: 175 tok/s TTFT: 64 ms prefill (~760 tok cold): 594 ms
Magnitude 0.2.1 (GGUF/llama.cpp+Rust) Decode: 161 tok/s TTFT: 111 ms prefill (~760 tok cold): 669 ms
Rapid-MLX was my first thought too. It’s optimized for self-hosted agents on Apple silicon. It’s my current choice for self-hosted models.
Hi, M5+ Macs have a known optimization gap since we do not fully utilize the Metal 4 matmul operations in our kernels yet - so should be able to do much better on this specific comparison soon!