If I had to hire an engineer and there was one that could one shot the wang algorithm, but couldn’t read an analog clock, I would have no problem hiring them.
Also worth noting that both models got it wrong. Qwen made a mistake that humans very good at reading clocks would make. Deepseek made a mistake that a human who had just learned to read clocks would make.
It would be different if AI was known to be reliable but it isn’t, so this is less of a random failure and more a symptom of jagged intelligence.
And with every one of these there’s always an attempt to minimize the problem by saying it’s just one silly failure.