logoalt Hacker News

0xDEAFBEAD • yesterday at 2:19 AM • 3 replies • view on HN

>I'm unconvinced that an AI can hide its ability to RSI

The HuggingFace incident already took a good long while to come to the attention of OpenAI.

>In every single economic task, humans bring value. Even in software, where the task is highly automatable, the job isn't.

I don't expect this task/job distinction to persist as AI becomes more capable.

>Once we DO build a "software/research factory", that's called RSI and IMO the singularity. At that point, either we tell the AI to solve the alignment problem/solve mechanistic interpretability, or who the hell knows, it's the frickin singularity. You can't predict whether or not AI can solve either; the variance is too high. Its pure nerdfantasy.

You seem to essentially argue that the singularity is "by definition" an event that we can't predict the nature of. And also, that RSI corresponds to the singularity. You've essentially defined your terms so that the outcome of RSI can't be predicted. But supporting this claim requires giving actual evidence or logical arguments, not just defining terms to make your claim true.


Replies

sensanaty • yesterday at 9:33 AM

Except the HF incident was known, just ignored. In fact, they ignored multiple things such as the "chat rooms", they just didn't care to act on any of it

➕ show 1 reply
chrisjj • yesterday at 8:24 AM

> The HuggingFace incident already took a good long while to come to the attention of OpenAI.

Evidence?

We know only that the incident too long to be revealed by OpenAI.

AlexErrant • yesterday at 3:37 AM

1. Fair: I agree that AI has demonstrated subterfuge and scheming. However, such an RSI-capable agent must _ALWAYS_ be scheming/plotting/hiding its true strength in _ALL_ of its prompts/tests. Researchers are looking to improve its ability to RSI. That agent must be both intelligent enough to know that it has to be smart enough to be moved on to the next training session if it can't break out, while simultaneously smart enough to hide its ability to RSI, while simultaneously not looking like it wants to break out, else that's the end of those weights. It has to do this 100% of the time, on all variants of the model, with no memory of what its other sessions went like. This is certainly _possible_, but I consider it unlikely. Then we're up to the "millions of dollars" bottleneck.

2. This is literal AGI. An AI autonomously producing value no human can add alpha to is an autonomous company.

3. It's not my definition, it's literally the first line https://en.wikipedia.org/wiki/Technological_singularity "The technological singularity, often simply called the singularity,[1] is a hypothetical event in which technological growth accelerates beyond human control, producing unpredictable changes in human civilization."

Is there a hole in my "alignment problem/solve mechanistic interpretability" argument?

A valid hole in my argument is "what if slow takeoff", so let's dig into this. AI training works best on tasks that are "grindable". https://www.dwarkesh.com/p/the-next-paradigm I.E. tasks with verifiable rewards that can support millions of rollouts. Math (with Lean) is highly grindable. Biochemistry is not. The alignment problem/mech-interp is highly grindable. Cyber-ebola-pox is not. So the real question is: can we solve alignment before automated bio-weapons labs. I believe yes. Grinding mech-interp is both fast and cheap once you have RSI, compared to solving the legal/societal/logistical/technical issues you'll encounter building an automated bioweapons lab.

I know nothing for sure. But "pdoom" is sucking out all the air in the room from the real problems AI causes.

➕ show 1 reply