Model is suspiciously fast and has a low reported output token count (using via OpenRouter's Chat), both of which aren't representative of models from the big Chinese labs. Odd.
> suspiciously fast
They're reporting ~30tps, that's about in line with many medium sized models served by Chinese providers
> Model is suspiciously fast
> aren't representative of models from the big Chinese labs
There were reports that China has let Nvidia's chips through, so this might be it. Testing both the chip and infrastructure.
GLM-5.3 is one of the faster models, at least according to artificialanalysis - openAI and Anthropic are the slowest.