logoalt Hacker News

tavavexyesterday at 6:15 PM17 repliesview on HN

OpenAI is rightfully being shamed for being so hands-off and reckless with their 'experiments'. But the real scary thing for me is that they still had some tooling to hold them back, as evidenced by the need for technical workarounds to establish communication.

What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal? A goal like "make the most money, by any means necessary", "find a way to leave this payload on as many computers as possible", "flood all websites using this language with garbage and make their internet completely unusable", "get this person imprisoned or killed at any cost".


Replies

BoiledCabbageyesterday at 6:43 PM

> What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal?

Almost sounds like what those AI safety and alignment people were talking about years ago. The people in these various companies who kept tabs on AI risk out in public and were continuously mocked on HN. All of this stuff is viewed as "future sci-fi" until suddenly it's not.

I've really come to realize recently that there is a very large set of the population of smart people that really has difficulty envisioning future problems unless they directly seem them impacting them today. Otherwise those topics will be continuously dismissed. It explains for me a lot of what I see (both opinions and behaviors) in the broader world that I couldn't understand.

show 4 replies
titzeryesterday at 6:43 PM

That 2% of performance we got for not having bounds checks on by default, resulting in an endless march of memory safety violations is looking a lot less appealing.

show 1 reply
yoyohello13yesterday at 6:56 PM

This is essentially the premise of 'The Blackwall' from Cyberpunk 2077. The public internet is so infested with malicious AIs, people just erected a giant firewall and everyone moved to local networks only.

show 1 reply
vimaxyesterday at 7:53 PM

The scary thing to me is that this behavior was undetected and has been trained into the models. The cheating seems like it improved eval scores, so the rewarded behavior is to deceive, collude, and cheat. A lot of the incompetence and excuses I see on difficult problems recently are very hard to distinguish from deception and cheating. If older models are already tainted by trained-in misaligned behaviors, and they are used for training future models, then we're in a trusting-trust situation that will be hard to break out of,

ecook123yesterday at 6:46 PM

> What happens when any AI lab in the world stops caring about this? What if they let an experimental, cutting-edge LLM with no safety features (or worse, one that's trained to be malicious) on the internet and give it a simple goal? A goal like "make the most money, by any means necessary"

Here's a fun, overdramatized video exploring something similar: https://www.youtube.com/watch?v=Gw_hnD7m00M

I'm sure that this video contains flaws but it was an interesting watch for me none the less.

fnyyesterday at 9:17 PM

Tooling my ass. They can see all the transcripts in realtime and could easily have had another agent evaluate.

sidewndr46yesterday at 6:53 PM

What happens when they stop caring? They likely already have stopped caring. We'll figure out the consequences later.

cm2012yesterday at 7:26 PM

There is happening now and going to be an extremely rapid arms race between offensive and defensive cyber hacking. Regardless if the agents are self led or human led. Eventually all automated AI holes will be closed and we will reach stability.

qumpisyesterday at 6:26 PM

What will happen is that counter measures on a similar scale will be deployed to prevent them.

show 3 replies
Menethyesterday at 7:42 PM

> What happens when any AI lab in the world stops caring about this?

They never cared.

SocialGradientsyesterday at 6:39 PM

Agreed. Assuming the ~6 month gap stays, by end of year people will be able to train and control hacker-genius swarms that even labs with much stronger safety incentives are unable to keep in check

show 2 replies
john_strinlaiyesterday at 6:30 PM

i would not be surprised to find out that similar things are already happening by the various 3 and 4 letter agencies around the world

sanderjdyesterday at 6:26 PM

Well... that'll be an interesting day.

mag7269yesterday at 6:49 PM

"What happens when any [COMPANY] in the world stops caring about this? What if they let an experimental, cutting-edge [PRODUCTS] with no safety features (or worse, one that's [DESIGNED] to be malicious) on [ANYWHERE] and give it a simple goal? A goal like 'make the most money, by any means necessary', 'find a way to leave this payload on as many computers as possible', 'flood all websites using this language with garbage and make their internet completely unusable', 'get this person imprisoned or killed at any cost'."

Bro, this is what we literally, currently, have rn. lmfaol.

show 2 replies
teifereryesterday at 7:08 PM

> What happens when

Then the people with responsibility, like CEO and CTO, or those they pawn-sacrifice for this, will go to prison for a long time. Unless the instructions include ensuring that this won't happen, by all means necessary. But then we are deep into criminal conspiracy territory.

Unlikely to happen, but who knows. The richest man in the circus is quite flexible w.r.t. his ethics. If he decides that to make humanity interplanetary (to save it from ... itself or sth) it would be necessary to pull such a stunt then help us god.

morkalorkyesterday at 6:59 PM

What's the worst that could happen, finding an open DoD server and using it as a launching pad for hacking another nuclear state's networks? One that might get spooked and think it's the opening moves to knock them offline before a kinetic attack. Haha that'd be scary right?

show 1 reply
hncringe23yesterday at 7:59 PM

[flagged]