I think M1 through M3 were compute bottlenecked in prompt processing (hence the very large gap between M3 and M5, in this page's benchmarks, that's not explained by memory bandwidth alone).
For generation speed in isolation, yes.
The M5 generation added tensor instructions to the GPU cores.
The M5 generation added tensor instructions to the GPU cores.