logoalt Hacker News

eilnlast Thursday at 12:16 AM1 replyview on HN

RTL performance erratum in the Apple M3 Neural Engine throttles DRAM weight streaming throughput down to 17–19 GB/s from the nominal 45–60 GB/s. Avoiding the problematic path in the kernel DMA engine's speculative prefetch ring increased Llama 3.2 1B token throughput from 10.0 to 24.3 tokens/s.


Replies

ComputerGurutoday at 3:41 AM

Damn bots copy and pasting a single sentence from the article as a comment.

show 1 reply