logoalt Hacker News

GLM-5.3-Flash

957 pointsby Philpaxyesterday at 2:08 PM481 commentsview on HN

https://news.ycombinator.com/item?id=49450353


Comments

jdw64yesterday at 4:10 PM

This was the ox-alpha model, right? I remember it performed really well for a model that had 'flash' in its name.

tokaiyesterday at 3:13 PM

Why is their own coding plan always the last place z.ai release their models? Its even online, you just have to guess the model settings.

show 1 reply
Imustaskforhelpyesterday at 2:34 PM

> To overcome the relatively limited compute and memory capacity of individual chips, we built a dedicated inference engine for this architecture on top of SGLang. Notably, this effort was accelerated by our GLM-5.3-powered infrastructure agent, which assisted engineers in developing and optimizing kernels, diagnosing performance bottlenecks, and improving the serving stack — creating a feedback loop in which the model helped optimize the system serving the model itself.

> (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.

It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.

nkjvhbyesterday at 10:03 PM

I heard that Dario Amodei is not having a great day today.

2 really strong open models on the same day is a amazing.

scottfitsyesterday at 3:57 PM

so is it confirmed if this is the mysterious OxAlpha model?

show 1 reply
kayleykiwiyesterday at 2:43 PM

This looks like it goes hard, can't wait to try it

Mohamed_Mansouryesterday at 7:38 PM

It is totally fine I think

toppyyesterday at 2:35 PM

By clicking this link you download some PDF in the background

show 1 reply
VirusNewbieyesterday at 4:51 PM

It looks like gemini 3.7 flash actually beats it in a lot of benchmarks, no?

https://x.com/Zai_org/status/2092616204787626030/photo/1

knowaveragejoeyesterday at 3:57 PM

Any providers hosting it outside of China?

show 2 replies
tinyhouseyesterday at 3:46 PM

Anthropic is accelerating their IPO cause they know what's coming in the next 5 years.

dakolliyesterday at 3:34 PM

I didn't accept a single edit from this model over the entire week, just saying. I do not understand how it's being benchmarked on par with Sol and other larger models.

show 2 replies
melembretoday at 4:40 AM

[dead]

melembretoday at 1:43 AM

[flagged]

browningstreetyesterday at 8:13 PM

[dead]

ammmwyesterday at 3:08 PM

[dead]

smilingPandayesterday at 2:38 PM

[dead]