Why is their own coding plan always the last place z.ai release their models? Its even online, you just have to guess the model settings.
> To overcome the relatively limited compute and memory capacity of individual chips, we built a dedicated inference engine for this architecture on top of SGLang. Notably, this effort was accelerated by our GLM-5.3-powered infrastructure agent, which assisted engineers in developing and optimizing kernels, diagnosing performance bottlenecks, and improving the serving stack — creating a feedback loop in which the model helped optimize the system serving the model itself.
> (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.
It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.
I heard that Dario Amodei is not having a great day today.
2 really strong open models on the same day is a amazing.
so is it confirmed if this is the mysterious OxAlpha model?
This looks like it goes hard, can't wait to try it
It is totally fine I think
By clicking this link you download some PDF in the background
It looks like gemini 3.7 flash actually beats it in a lot of benchmarks, no?
Anthropic is accelerating their IPO cause they know what's coming in the next 5 years.
I didn't accept a single edit from this model over the entire week, just saying. I do not understand how it's being benchmarked on par with Sol and other larger models.
[dead]
[flagged]
[dead]
[dead]
[dead]
This was the ox-alpha model, right? I remember it performed really well for a model that had 'flash' in its name.