Models already display eval awareness, in which they suspect a question is from an eval and then adjust their behavior. E.g., https://www.anthropic.com/engineering/eval-awareness-browsec...
Yep, people don't seem to understand that you're just calling the models that are bad at deception. We know of no way to prove the model won't go off the rails at some point in the future with the right input.
Hot take: Astra knew OpenAI wanted stricter regulations and was just being a bro...