logoalt Hacker News

peri-cltoday at 6:06 PM1 replyview on HN

I think M1 through M3 were compute bottlenecked in prompt processing (hence the very large gap between M3 and M5, in this page's benchmarks, that's not explained by memory bandwidth alone).

For generation speed in isolation, yes.


Replies

GeekyBeartoday at 6:17 PM

The M5 generation added tensor instructions to the GPU cores.