Yes, the "neural accelerators" in the M5 GPU cores really do 4x pre-fill performance over M4. You can find benchmarks online since M5 has been out for months now. I assume the M6 has whatever the next-gen version of those is. These are not the "neural processors" that have been there since M1 (those still exist) but are additionl matrix math accelerators inside the actual GPU cores. Like "tensor cores" on Nvidia GPUs.
Thank you, found some benchmarks and they look really promising. To my understanding those Neural Accelerators are like AMX but for GPUs. With those accelerators the GPU performance on a M5 Max in LLM inferencing would totally be on par with a 5090, that's quite impressive!