Hey, yeah this is a known issue on M5+ macs. We are working on a patch so that our kernels use that hardware acceleration path. This should make prefill faster than MLX-based engines and boost decode a bit more for that hardware!