This line made me think by 'normal coding work' the author means doing something they don't understand well enough to be able to distinguish the models' output.
I've been trying Chinese models, like GLM 5.2, as substitutes for Claude or GPT5.X, and my experience is that they underperformed them in real world metrics like prompt adherence, hallucination or code quality, even though they benchmark better.
Steve Yegge calls this is the "discernment horizon" - https://steve-yegge.medium.com/the-flat-curve-society-36c8b0...