Very informative article. I was hazy on where NVLink sat in the local inference stack. This makes it very clear: to adapt consumer hardware that lacks multiple full x16 PCIe lanes.
I badly miss the days when I could buy a GPU and then three years later buy the same GPU for a hundred bucks and boost my framerate by 70%.
Thanks
Great piece. Was wondering whether to get an NVLink bridge for $500+ for my two A4500s for inference. This pretty much confirms it only really rescues you if you have nonsensical PCI-E topology.