logoalt Hacker News

Casteilyesterday at 2:40 PM2 repliesview on HN

I don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf.

qwen3.5:122b-a10b is significantly faster at around 60-65.


Replies

smcleodyesterday at 8:47 PM

No magic, just oMLX with MTP. You can look through the speed the community is getting here: https://omlx.ai/benchmarks/performance?model=qwen3.8&chip=&c...

syntaxingyesterday at 2:54 PM

With MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further

show 1 reply