logoalt Hacker News

francisjp • last Wednesday at 10:31 PM • 1 reply • view on HN

To OP: great work on the release! I am generally interested in this kind of optimization work.

Related to the post above: Similar results here M5 Max running Qwen3.8 UD-Q6-K-XL with zlab’s Dflash2 as the drafter.

Both prefill and decode are roughly 2x faster when served from llama.cpp (b10853 or newer) than magnitude 0.2.1.


Replies

anerli • last Wednesday at 10:45 PM

Thanks for pointing this out. I think we have a gap here where we may not be fully leveraging the new matmul operations available on M5+ chips, so will work that into our kernels soon and benchmark on M5 hardware.

➕ show 1 reply