logoalt Hacker News

mrtsepelev • today at 10:08 AM • 1 reply • view on HN

Congrats on launch! Tried it on the gemma-4-26b-qat-4bit model. Was indeed faster on token generation then on oMLX (82.8 tok/s vs 76.5 tok/s), but the prefill time was ~2.6x slower (709 tok/s vs 1843 tok/s). Don’t use any acceleration on the oMLX. Macbook M5 Pro, 48 gb


Replies

anerli • today at 6:11 PM

Hey, yeah this is a known issue on M5+ macs. We are working on a patch so that our kernels use that hardware acceleration path. This should make prefill faster than MLX-based engines and boost decode a bit more for that hardware!