logoalt Hacker News

angoragoatstoday at 10:55 AM0 repliesview on HN

Have you ever actually set up a multi-GPU system for inference? Based on my experience you are drastically overstating the problem. Both tensor and pipeline parallelism (without NVLink) produce a machine which is faster than any Mac on the planet, which is what we’re discussing here. Yes, each has pros and cons, and neither scales perfectly linearly. But it works great regardless.