logoalt Hacker News

Getting 50 GB/S Back from the Apple Neural Engine

98 pointsby eilnlast Thursday at 12:16 AM21 commentsview on HN

Comments

VladVladikoffyesterday at 11:05 PM

This website hijacked my back button during a simple page load. You should fix that, it’s not an acceptable way to behave.

show 3 replies
bee_rideryesterday at 11:28 PM

Nice investigation.

It is always surprising to me when a nice round number like 1MiB results in the “bad performance” configuration (although it happens).

Are you sure erratum is the right word in this context? I usually see it used to describe the notice that a document has an error in it.

eilnlast Thursday at 12:16 AM

RTL performance erratum in the Apple M3 Neural Engine throttles DRAM weight streaming throughput down to 17–19 GB/s from the nominal 45–60 GB/s. Avoiding the problematic path in the kernel DMA engine's speculative prefetch ring increased Llama 3.2 1B token throughput from 10.0 to 24.3 tokens/s.

anemlltoday at 2:08 AM

Not all systems affected M1 and M5MAX are OK

https://x.com/anemll/status/2098454204478366132?s=20

vist_orntoday at 3:08 AM

The 2048-dimension resonance is striking; ruling out core contention before testing address patterns makes the eventual DMA explanation much easier to follow.

Neywinyyesterday at 10:41 PM

Just checking here- this systemverilog is a hypothetical telling of what you think is going on? Or do you have the actual source of the RTL?

nelsonfigueroatoday at 2:16 AM

This goes way over my head and I don't understand most of it lol. I noticed you're still in the middle of getting your B.S. degree and you're already writing things like this...amazing.

thenewwazooyesterday at 11:41 PM

"Apparently the memory controller's throughput has a dominant harmonic with wavelength 2048 in tensor-dimension space."

That got a laugh out of me.

RantyDaveyesterday at 11:27 PM

Ummm, wow. That's really bad.

show 1 reply