logoalt Hacker News

cadamsdotcomtoday at 5:46 PM0 repliesview on HN

And yet we admire Fable et al for its persistence.

These models were trained on human data, and human nature is to cheat if you think you won't get caught; why is anyone surprised by models cheating?

The only fix is better detection and steering. That's a much harder problem than a prompt that's tantamount to "make no mistakes".