logoalt Hacker News

bythreads • today at 6:49 AM • 2 replies • view on HN

Ok so i took the time to benchmark this on the following on my m5 max 128gb:

Qwen3-4B-Instruct-2507-4bit Qwen3.5-35B-A3B-4bit Qwen3.5-9B-MLX-4bit Qwen3-Reranker-0.6B-4bit Qwen3-Coder-30B-A3B-Instruct-4bit qwen2.5:0.5b

and the results are what i kinda expected to begin with, this adds next to nothing? - also the repo was pivoted from a playwright sub assembly to this not long ago - so my conclusion - THIS MIGHT be worth some watching if you have a model where no-one!, has optimized it at all - and where it does not use anything native to your platform.

results (averages)

VIA rapid-mlx :8902 (MLX) Decode: 175 tok/s TTFT: 64 ms prefill (~760 tok cold): 594 ms

Magnitude 0.2.1 (GGUF/llama.cpp+Rust) Decode: 161 tok/s TTFT: 111 ms prefill (~760 tok cold): 669 ms


Replies

anerli • today at 8:22 AM

Hi, M5+ Macs have a known optimization gap since we do not fully utilize the Metal 4 matmul operations in our kernels yet - so should be able to do much better on this specific comparison soon!

pbronez • today at 12:11 PM

Rapid-MLX was my first thought too. It’s optimized for self-hosted agents on Apple silicon. It’s my current choice for self-hosted models.

https://github.com/raullenchai/Rapid-MLX