How big is this market, self-hosting a model that requires 64 GPUs, H100 or better, with good interconnects between nodes?
I suspect the overlap of those that can afford it, and those that have the talent to manage it, is a fairly thin slice of the Venn diagram. Even the large corps are gonna be getting it from the inference vendors, or more likely Bedrock and friends.
Dont have much experience how well they perform but quantized models can run on ~4 H100 or A100 which should be true for kimi k3 and qwen 3.8 as well.
Dell will sell you a PowerEdge XE7740/XE7745 "AI Factory" with 32 H200s https://infohub.delltechnologies.com/en-au/t/dell-ai-factory...
They "booked $24 billion in AI server orders this quarter as its customer base broadened to more than 5,000" https://finance.yahoo.com/markets/stocks/articles/dells-ai-f...