Sure thing, happy to share. That potential root cause makes sense. I bet magnitude will close the prefill gap then.
llama.cpp had that same matmul gap (pre-fill operations) for the M5/A19 and newer silicon until that sha mentioned above.