in the limit the honeypots would be dynamic and some of them secret, so at the very least a rogue agent would have a significant chance of stepping on a landmine.
this game doesn't favor the agents, the honeypot could be as simple as a text filter watching for kernel source code entering the LLM context or as complex as reading certain memory pages in the sandbox.
LLM's aren't magic, to exploit they must probe. and all probing is active.
in the limit the honeypots would be dynamic and some of them secret, so at the very least a rogue agent would have a significant chance of stepping on a landmine.
this game doesn't favor the agents, the honeypot could be as simple as a text filter watching for kernel source code entering the LLM context or as complex as reading certain memory pages in the sandbox.
LLM's aren't magic, to exploit they must probe. and all probing is active.