> Mac studio wins in memory capacity, price, perf/watt and value.
[citation needed]. I have personally specced out and built an nvidia GPU-based machine which after some optimization, handily beat the Mac Studio in terms of tokens/watt for LLM inference with most models. This was in the M2 Ultra era, and I haven't run the numbers for the later generations, but nvidia's cards have gotten faster just as Apple's CPUs/GPUs have, so I would guess that it's still possible to do.
> RTX 6000 wins in performance, if your model can fit into the VRAM.
"if your model can fit into the VRAM" can be true for the Mac as well.
> There are very obvious and clear advantages to a Mac Studio.
There are certain advantages for sure, depending on your use case. They may _seem_ to be obvious, but as evidenced above, I believe that many people overestimate the Mac's superiority on the metrics you cite when comparing a Mac vs. a dedicated GPU for LLM inference.
[citation needed].
No need. You can infer the logic with this line I wrote: RTX 6000 wins in performance, if your model can fit into the VRAM.
I'm not sure what the controversy is here.
> "if your model can fit into the VRAM" can be true for the Mac as well.
It is much more likely for your model to fit in large unified memory of a Mac than the smaller more limited memory of a GPU. Even going with two 5090s, you now have to shard your model and that is a PITA.
But it turns out that MoE is the solution both for running models on macs of limited computer power means (not as fast as GPUs), and on multiple GPUs that require sharding the model.