Ox Alpha is a smaller model and it was running very slowly. Chinese AI accelerators are coming along, but nVidia’s lead is huge.
It was being served for free. They were almost certainly being overloaded.
Ox Alpha was also serving 10T+ tokens a day for free.
When it first launched on OpenRouter I was getting nearly 70 Tokens/second.
Has there been any confirmation about what that model even is?
Edit: Ah:
> This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash.
> and it was running very slowly
... I'm at a loss for words here. It was being served for free. To the entire world.
Lead doesn't really matter anymore. I just ported a very old cuda library to rocm, so it can be run on MI300s. 2 years ago this would have been a nightmare. Today it was an afternoon.