Surely OpenAI could adjust their RL to sharply penalize cheating. They have access to the full traces including reasoning: it can’t be that hard to detect an attempt to find the solutions outside TVs space that is fair game for exploit attempts (making sure that a description of the valid targets is in the prompt).
For that matter, if a rollout breaks out of the sandbox, they should detect it, pause, and fix the bug.