logoalt Hacker News

spindump8930 • yesterday at 1:48 AM • 0 replies • view on HN

The problem (and contrast with other approaches) is that mat muls requires synchronization. Arranging your networking and training structure to maximize compute and minimize communication is the main craft of ML training infra folks. In your example, yes you can compute layers on different machines (i.e. Tensor Parallelism), but you must be very careful in how you arrange it.