logoalt Hacker News

dada216today at 10:07 AM2 repliesview on HN

We built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system.


Replies

freakynittoday at 10:46 AM

Most of the people had kinda guessed this when they decided to provide 100 trillion tokens for free.

axiosgunnartoday at 10:13 AM

[dead]