We built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system.
Most of the people had kinda guessed this when they decided to provide 100 trillion tokens for free.
[dead]
Most of the people had kinda guessed this when they decided to provide 100 trillion tokens for free.