I’ve run the numbers on this, and an optimized nVidia build can meet or beat Apple platforms in terms of tokens per watt, which is the most important efficiency metric if what you care about is using the least amount of energy to generate a given response.
Yes, the Mac might get lukewarm, but it will take 2-3+ times longer to do the same task.
Yes I agree that at full speed with parallel workloads, the tokens per watt are better using Nvidia GPU servers.
Also while the pre fill performance sucks, the Mac isn’t that slow and can use much better model compared to a similarly priced Nvidia workstation so it’s not really taking much longer in practice.
It’s taking longer than in the cloud for sure. At least a local computer uses the local energy grid that is pretty clean and not fossil energy.