logoalt Hacker News

stefsyesterday at 8:28 PM1 replyview on HN

that's not a solution, that's just an impediment - at least if the agent can keep persistent memory about its state between attempts.


Replies

teravoryesterday at 8:38 PM

in the limit the honeypots would be dynamic and some of them secret, so at the very least a rogue agent would have a significant chance of stepping on a landmine.

this game doesn't favor the agents, the honeypot could be as simple as a text filter watching for kernel source code entering the LLM context or as complex as reading certain memory pages in the sandbox.

LLM's aren't magic, to exploit they must probe. and all probing is active.