Thanks for pointing this out. I think we have a gap here where we may not be fully leveraging the new matmul operations available on M5+ chips, so will work that into our kernels soon and benchmark on M5 hardware.
Sure thing, happy to share. That potential root cause makes sense. I bet magnitude will close the prefill gap then.
llama.cpp had that same matmul gap (pre-fill operations) for the M5/A19 and newer silicon until that sha mentioned above.
Sure thing, happy to share. That potential root cause makes sense. I bet magnitude will close the prefill gap then.
llama.cpp had that same matmul gap (pre-fill operations) for the M5/A19 and newer silicon until that sha mentioned above.