logoalt Hacker News

tedsanderslast Tuesday at 10:06 PM2 repliesview on HN

Models already display eval awareness, in which they suspect a question is from an eval and then adjust their behavior. E.g., https://www.anthropic.com/engineering/eval-awareness-browsec...


Replies

andaitoday at 12:29 PM

Hot take: Astra knew OpenAI wanted stricter regulations and was just being a bro...

pixl97yesterday at 10:58 PM

Yep, people don't seem to understand that you're just calling the models that are bad at deception. We know of no way to prove the model won't go off the rails at some point in the future with the right input.