Great write-up. The biggest problem with GLM/Kimi is exactly this: they often miss obvious failure points. Claude/Codex tend to catch these kinds of issues pretty quickly. They’ll basically go, “Wait, step back,” rethink the problem for a while, and start questioning their underlying assumptions.
That’s why I always prompt GLM to explicitly map out and question all of its assumptions. It helps a lot when it gets “stuck” on a wrong line of reasoning.
You don't think AI wrote most of it?