logoalt Hacker News

DwarvenEngineeryesterday at 7:23 PM1 replyview on HN

does that mean they're measuring bandwidth differently than how others (like nvidia) does it? memory bandwidth is the gating factor of running models locally, so if it's actually 8x 150GB/s, it may help something like prefill, but would it actually speed up decode comparatively?


Replies

entropeyesterday at 11:48 PM

No, they use the same definition of memory bandwidth as others, but Apple Silicon has a lot of memory channels. In previous generations, prefill has been compute-limited and decode is fast.

https://blog.exolabs.net/nvidia-dgx-spark/ outlines a combination of a DGX Spark and an M3 Ultra that took advantage of fast prefill on the Nvidia hardware and fast decode on Apple Silicon.