And yet we admire Fable et al for its persistence.
These models were trained on human data, and human nature is to cheat if you think you won't get caught; why is anyone surprised by models cheating?
The only fix is better detection and steering. That's a much harder problem than a prompt that's tantamount to "make no mistakes".