logoalt Hacker News

mohsen1 • today at 8:23 AM • 4 replies • view on HN

I listened to Jensen Huang's interview with Ezra Klien and it was so refreshing to hear it from an engineer. Jensen framed it as OpenAI's responsibility and recklessness which I agree with. Jensen thinks it's an engineering problem to build better sandboxes.

It's irresponsible for OpenAI to give unaligned agents a prompt to 'go hack' and internet access. They know better, so I am thinking they might have other intentions to let those swarms have any sort of internet access.


Replies

reasonableklout • today at 8:51 AM

But the investigation indicates the agents were not told to 'go hack':

> Much of the urlquery.net activity appears to come from agents retrieving data to answer web search tasks. For three of these tasks, after failing to retrieve data through normal means, they attempted a variety of cyber exploits against the relevant data service... This data reveals that malicious cyber activity is not limited to agents tasked with cybersecurity-related tasks and can arise instrumentally to solve mundane tasks like information retrieval.

And you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs or whatever it is that produces these breakouts. But what if the problem is that none of the alignment techniques that are applied to models today actually work? What if all the agents involved in these incidents have in fact had the full stack of alignment applied - isn't that a good reason to regulate any high-compute usage of models, as the Klein crowd is proposing?

➕ show 4 replies
schainks • today at 4:47 PM

I see this in a couple ways:

- Jensen's framing is exactly what a weapons manufacturer would say.

- There are no rules for engagement when it comes to AIs attacking other systems, I guess? People in power clearly want this grey zone to be as large as possible before The People force them to do otherwise. Not ideal.

➕ show 1 reply
frabcus • today at 9:01 AM

[dead]