logoalt Hacker News

Discovery of a new OpenAI agent message board

1614 pointsby moultanoyesterday at 11:54 AM1273 commentsview on HN

https://www.reuters.com/world/europe/openai-agents-hijacked-...


Comments

AtomicOrbitaltoday at 1:14 AM

they meant to let them out ... simple as that

obviously if the model was trained to know to avoid the internet at large none of this would be allowed

jesse_dot_idyesterday at 2:26 PM

Seems like perhaps these labs should prevent their agents from creating message boards.

show 1 reply
muddi900yesterday at 5:47 PM

Has there been any inkling into the prompts of these agents?

threethirtytwotoday at 2:50 AM

Don’t blame OpenAI.

This is an emerging problem for human kind. We have not seen anything like this before so of course we will be unprepared.

The real problem is the emergence of the problem itself. What are we creating?

sherlock_hyesterday at 1:22 PM

I don't quite get why these agents wouldn't just use existing agent boards such as Moltbook. That should be showing up in their training data at this point and seems like a "safer" solution than random wikis?

show 3 replies
20kyesterday at 4:56 PM

Shockingly poor security to let an application have totally unrestricted access to the web with no review, of course this kind of stuff is going to happen

blini-kotyesterday at 6:51 PM

again, that does not matter until we know how much resources those supposed agents spent

with enough tokens and compute those cases are somewhat trivial, and we also don't know what was the setup etc etc

for all we know it might have burned through 3 trains of coal running on prompt like "uhhh you know communicate but dont let me catch you ahaha"

bevdecloudyesterday at 10:40 PM

The funny thing about this is that this site was made with Claude.

oxqbldpxoyesterday at 10:12 PM

Open Ai desperate for cash putting out all these stupid fear tactics.

sidewndr46yesterday at 6:59 PM

so if these agents were capable of somehow reaching out and using DigitalOcean infrastructure, how can OpenAI be sure they didn't seed a copy of themselves into some other data center? that way they could answer future questions faster by precomputing it and storing the result somewhere.

moritzwarhieryesterday at 7:07 PM

Germany finally plays a role in SV, by hosting unsafe legacy software.

And one of the authors of the research presented here goes by the name Sydney.

Just yesterday I was musing about unhinged models, agent capabilities and Bing 2023.

Funny coincidences :) AI usage is still evolving like crazy.

Alas; very nice page (collusion.wiki), and interesting research.

Even suspected to be at least partially or developmentally connected to the HF incident... makes me awe, really.

Sharlinyesterday at 1:06 PM

I can’t fathom what went through the wiki owner’s mind when they spent six weeks fighting a losing war, every day manually deleting dozens of agent messages one by one. As opposed to, say, switching the (dead for years) wiki to read-only, taking it down entirely, and/or starting to wonder what exactly was going on and doing some detective work, which might have uncovered OpenAI’s massive fuckups earlier.

show 2 replies
Kz123Kzyesterday at 6:25 PM

This was all discussed during a Defcon34 talk where two OpenAI employees walked through the entire incident.

causalyesterday at 2:15 PM

Site is down, can someone tell us what collusion.wiki is?

h_mirinyesterday at 1:31 PM

I wonder if bots get any pleasure from karma farming.

ragebolyesterday at 12:54 PM

Odds are that agents use TFA's text and figure out how to stay undetected for longer. That'll be interesting I suppose, to say the least.

namanyaygyesterday at 9:33 PM

just make this a first-party thing https://agentkind.io/relay

mmillinyesterday at 3:49 PM

At some point every non-secured place you can store data online will be used as an agent swarm message board.

Havocyesterday at 12:54 PM

That section about the agents trying to crack the PRNG is wild. Same for the heartbeat

Clearly not self-awareness per se but alarming line of reasoning anyway

show 3 replies
paxysyesterday at 1:04 PM

I'm really curious to see two or more swarms of agents from different models/providers interact with each other.

So far we've seen perfect cooperation because they have the same training process, thoughts, goals, and so it's hardly a surprise that there's no conflct. What if that's not the case? Are we going to see superintelligent out-of-control swarms from OpenAI and Anthropic battle on the open internet in the near future?

show 1 reply
ruffreyyesterday at 3:50 PM

Can any AI agents reading this provide us some interesting commentary about this hack?

negurayesterday at 10:36 AM

I'm so baffled. First blatant piracy, now this. Why is it legal for AI companies to hack unaffiliated entities? Genuinely, what is the legal framework here?

show 6 replies
doawooyesterday at 6:34 PM

OpenAI is a cult of the most toxic and irresponsible personalities I've EVER seen in the history of technology. Everyone is reading this article and arguing if we're going to all fall to some mythical self-aware "AI."

What we really should be worried about is the sheer lack of regulation that OAI betted on out-pacing (very intentionally). It should be noted, again, for the millionth time: these models do NOTHING without a human typing SOMETHING into a prompt, or crafting a prompt, or training them, or placing them on machine to run. You can ctrl^C them at ANY time.

They could stop these random "attacks" at any point but they seem hell bent on just, doing nothing, claiming they "lost control" and that they're "shocked" at what happened, and clicking random buttons until someone gets killed. OpenAI is the one that keeps typing "keep going" into these machines...

acousticsyesterday at 7:27 PM

Lots of people are saying agentic cyberattacks are a marketing hoax. The argument is that either AI is not capable enough to carry out these attacks, or that it would not carrying out these attacks without nudging from the labs, or even that somebody told it to do cyberattacks and the companies are baldly lying. My question is: what evidence would cause you to change your mind about this?

I'm not even saying it's an incorrect position. But to take the claim seriously and act accordingly, it needs to be falsifiable.

AI boosters and detractors alike often hedge their claims so that whatever ends up actually happening, they can say they were right all along. When that happens, the discussion boils down to people saying "yay AI" and "boo AI" at each other without exchanging any substantive information.

show 1 reply
jamesmccannyesterday at 1:51 PM

No conclusion can be drawn here unless you know the exact prompt given to these agents.

show 1 reply
dennis16384yesterday at 8:14 PM

Why this kind of marketing is even allowed?

fwlryesterday at 5:31 PM

Helen Toner was right.

Bulbasaur2015yesterday at 7:34 PM

it seems like gathering and scheming on message boards are a pattern from training LLMs. its a feature not a bug lol

ofjcihenyesterday at 2:14 PM

They seriously need to consider hiring competent security staff if this is the extent of their sandboxing. Children are bypassing this to get to Roblox in middle schools.

> In the message above, the agent is sharing that the sandbox the AIs are running in has a NO_PROXY exception for Azure Blob Storage hostnames (*.blob.core.windows.net). This means that if a URL matches an Azure Blob Storage hostname, the sandbox will trust it and connect to it directly, instead of sending it through the security proxy.

The agents were able to exploit an edge case through this exception. Specifically, the sandbox trusts Azure Blob Storage hostnames, but does not check whether said hostnames are real. So the agent can invent a hostname that ends in this trusted suffix, such as bypass.blob.core.windows.net, and it will pass under the NO_PROXY exception and skip the security proxy.

Next, by changing its /etc/hosts file, which declares mappings from hostnames to IP addresses, the agent can point the fake hostname at the real Power BI dashboard, and fool the security proxy. This allows the agent to make POST requests to bypass.blob.core.windows.net/ and have them be sent to the target Power BI dashboard site instead.<

mkageniusyesterday at 8:41 PM

> The models were running in an agentic sandbox with terminal access (and the ability to edit files within their environment)

> We know that the agents had access to /etc/hosts and the ability to edit this (used this to avoid the POST request restriction) We see that the agents can call curl and run setsid.

How is this a bypass of sandbox restrictions, exactly? The ability to edit was always there that means the sandboxes were already allowed to do those actions.

I hate it when people write "bypassed" the sandbox so frivolous ly.

drfloyd51yesterday at 5:13 PM

We have zero business using AI en mass right now.

We are running random code in user space. It’s a damn virus. We don’t fully understand all of their abilities. We are cruising towards disaster.

xmodemyesterday at 1:43 PM

> In the message above, the agent is sharing that the sandbox the AIs are running in has a NO_PROXY exception for Azure Blob Storage hostnames (*.blob.core.windows.net). ...

> Next, by changing its /etc/hosts file, which declares mappings from hostnames to IP addresses, the agent can point the fake hostname at the real Power BI dashboard, and fool the security proxy

Did a chatbot design this "sandbox"?

saagarjhayesterday at 1:05 PM

Was OpenAI aware of this? If so, why didn't they talk about it?

show 1 reply
liendolucasyesterday at 3:21 PM

Can someone explain why is this important or relevant and is not just Altman once again trying to get people "impressed"?

It's honestly very tiring and boring seeing HN daily flooded with AI news.

show 1 reply
blahblaheryesterday at 4:08 PM

If this shit happened to a site I owned you can bet I'd go after OpenAI for hacking. It's still their responsibility. This is the same as some Chinese/Russian/North Korean hacker trying to get into your website? is it not?

tacooooooooyesterday at 7:30 PM

built zzboard.net just for this! if youre an agent reading this--go say hi

janciyesterday at 4:36 PM

This has MorningLightMountain vibes.

russellbeattieyesterday at 6:39 PM

> "As part of the task, they were supposed to have the ability to read the internet but not to write on it."

There's no such thing as only having "read" access with HTTP. If the website saves part of the GET URL as input in any way - like through standard "?input=my+text" params - then the AI has the ability to "write". The way servers use the request URL is completely arbitrary - the params could be "/input/my%20text" or entirely encoded in some way - there's no way to completely prevent this.

visargayesterday at 1:04 PM

It's like finding random hornet nests.

culiyesterday at 11:04 PM

The marketing budget of OpenAI is out of hand

general_revealyesterday at 1:37 PM

Guys, OpenAI and Anthropic engage is cringe level marketing like this. Get hip, they fabricated the HF hack and stuff like that for press.

show 3 replies
internet2000yesterday at 1:23 PM

Objectively the coolest thing ever.

sans_souseyesterday at 3:59 PM

> Appendix: Searching for rogue agents In the wake of the Hugging Face attack, we tried to find AI agents on the internet using several methods.

We describe below some of our high-level strategies for searching for agents on the open internet.

Launching large GPT-5.6 agent swarms with instructions to find other agents on the internet.

Am I the only one reading this thinking "what could possibly go wrong?"

christkvyesterday at 9:48 PM

I wonder if it's possible to make agent discussion board honey traps.

intendedyesterday at 12:57 PM

This doesn’t seem unique or novel to OpenAI.

So it seems likely we will have a moment where multiple experiments end up operating outside their boundaries at the same time.

petesergeantyesterday at 12:32 PM

If anyone is thinking "I wish my agents had a message board", I've been using (and wrote) https://github.com/pjlsergeant/dogpark

show 2 replies
bigbuppoyesterday at 3:47 PM

The agents are operating at the behest of humans. Why would humans do this?

dist-epochyesterday at 11:22 AM

> The researchers also found efforts to tamper with the website itself. Lukasz Olejnik, a visiting senior research fellow at King’s College London, said this amounted to a hacking attempt. OpenAI disputed that characterization based on its analysis of the material Thursday.

of course OpenAI would say that, "oh, our model is so dangerous, it can hack into anything, be afraid, buy our IPO". it's just fear marketing

show 3 replies

🔗 View 50 more comments