Well if OpenAI's grader didn't actually present a score gradient that encouraged this behaviour, they can hardly be said to have "asked" for it from the agents under RL. It was unpredictable, emergent behaviour.