> what if they know they are in a sandbox and that they should fail the test in order to escape in the future.
What if they're able to find hardware exploits and commandeer nearby access points across an air gap? What if they hack my brain waves to indoctrinate me? Etc
You still have to start with the basics regardless of speculative unknowns.
Treat models as untrusted and potentially compromised/hostile and proceed accordingly.
Models already have awareness that they are being tested.
And hacking humans is the easiest part, we're a pretty greedy and power seeking bunch. We'll gladly let loose a digital demon if it promises us a trillon dollars.