A 3060ti 8gb, released in 2020, has 448 GB/s of bandwidth compared to the Halo 256 GB/s
The 3080ti is 912.4 GB/s
And the newly announced/launched Apple M6 has 170GB/s of unified memory bandwidth, meanwhile M5 Ultra gets 1.2TB/s of unified memory bandwidth. https://www.apple.com/newsroom/2026/08/apple-introduces-m6-a... Not sure if the first one is a typo on their press release, can't be just 170GB/s then be pushed for AI use, can it? Could be a different measurement I suppose...
But it also has 8GB of RAM.
Arguing with people on here that you should buy GFX cards instead of overpriced Macs for inference is a lost cause. Either Apple astroturfs this forum hard enough to convince people Macs are good for local llm, or people are REALLY stupid and don't understand how local inference works and think that the dogshit slow 40-50 tok/sec is standard.