From a bandwidth perspective, the ultra is like 8 M5s fused together (@ 150GB/s), that's how it gets to the 1200.
Historically the Pro doubles the base, the Max doubles the Pro, and the Ultra doubles the Max.
If an M6 ultra were released today it would be 1.36TB/s.
does that mean they're measuring bandwidth differently than how others (like nvidia) does it? memory bandwidth is the gating factor of running models locally, so if it's actually 8x 150GB/s, it may help something like prefill, but would it actually speed up decode comparatively?