I've tested GLM 5.3 on the release day and Artificial Analysis is spot on. It's a really good model.
But my main takeaway was something else. I've used closed weight models for long enough that I've forgotten how good it feels to see reasoning tokens.
With GPT/Claude, you kind of hope that intent was captured well, that agent had all the information, all the tools it needed, because you won't see "hmmm it seems like nix flake isn't available here and I shouldn't install something globally" until it slopped out millions of tokens and wasted hundreds of dollars for 8 hours. With GLM and the likes, you just stop the disease right where it begins.
With GPT/Claude, hiding those from users to waste their tokens is a feature, not a limitation.
Yes, not necessary often but being able to stop something that is going off the rails is super useful. Especially if the root cause is prompt ambiguity - inject a clarification & it recovers