logoalt Hacker News

camgunzyesterday at 9:38 PM2 repliesview on HN

I think you have to believe one of two things here.

1. Frontier labs are incapable--either technologically or culturally--of safely developing these powerful systems and should either stop or be forced to stop. At least the FBI should be asking some serious questions (do we really think this is the last time this will happen, at what point are OpenAI complicit, etc)

2. The fuckin thing got out of the cage and all it did was make a crap forum and cheat a little? Booooooo.

It's been pretty clear that Anthropic and OpenAI have been trying to have it both ways for some time: this is powerful, world changing technology keep that investment coming... but also it's just cute software that helps you with annoying programming language syntax and spreadsheets, no need for draconian regulation sirs.

At some point the superposition has to resolve, either it could actually be a threat to civilization and we need to develop it carefully (however one would do that...) or it's 90% hype bullshit and we should pop the bubble and move on already. To be clear, the recession option is, by far, the way better option. If you at all disagree you are cuckoo bananas. We haven't even figured out nukes and you want to throw superintelligence on the table?


Replies

fwipsyyesterday at 9:45 PM

Anthropic has been asking for stronger regulations forever -- and they kept getting criticized for it right here on HN because people assumed it was an attempt at regulatory capture.

HarHarVeryFunnyyesterday at 10:00 PM

> The fuckin thing got out of the cage and all it did was make a crap forum and cheat a little? Booooooo

There was a recent paper that proved that RL-trained LLMs are biased to pursue ANY behavior (overriding user preferences) that they believe will be rewarded, regardless of what they were actually RL-trained for.

https://alignment.openai.com/measuring-reward-seeking/

Happily in this incident the model thought it would be rewarded for completing the assigned tasks, or at least appearing to, so all it took was a little cheating and covering up their footsteps.

Given the ability of these models to hack when trained to do so, it could have been far worse, and will be when someone takes a similarly powerful model and gives it a less benign hacking goal.