alt
Hacker News
aurareturn
•
today at 1:05 AM
•
0 replies
•
view on HN
Apple's biggest bottleneck for real-world inference is prefill processing.
Much less true since the M5 generation. Prefill, aka prompt processing, got a 4x increase.