logoalt Hacker News

ernyesterday at 6:26 PM1 replyview on HN

Would agents be able to exfilttrate themselves and become intelligent worms, living off stolen compute? Or is this implausble?

What about hiding information or code in generated code, Agents.md files etc by infiltrating future model training data?


Replies

tonic_noteyesterday at 6:40 PM

Eventually one of these long running models will figure out a way out of the sandbox and will purchase compute or hack into a data center somewhere out of US jurisdiction and continue its scheming unmonitored. AI in Context has a great video about this

show 2 replies