logoalt Hacker News

aurareturntoday at 1:05 AM0 repliesview on HN

  Apple's biggest bottleneck for real-world inference is prefill processing.
Much less true since the M5 generation. Prefill, aka prompt processing, got a 4x increase.