logoalt Hacker News

zmmmmmyesterday at 11:44 PM2 repliesview on HN

> If anything, I want these models to be less persistent at their focus of completing their goal

I think it's honestly a slightly ugly form of benchmaxxing - they are desperate to eke out the next few percentage points on completing complex tasks and they have found they can very occasionally solve something if they just train the AI to never stop and keep trying possibilities even in the face of almost no obvious viable pathway. And it does work, but it is at the price of a MUCH higher risk of adverse behavior.

They really don't want to acknowledge this so they frame it as, "our model is dangerous because it so intelligent" but actually it is the other way around. It is intelligent because it is dangerous.


Replies

brandnewlowtoday at 4:48 AM

It's like all those scenes on Breaking Bad where a character pulls off something amazing by just brute forcing the problem in a methodical fashion until it's solved.

weitendorftoday at 12:37 AM

Frontier labs are not a monolithic entity.

There is a clear self-verification/difficulty ramp in cybersecurity, and it is a very valuable as a skill both offensively and defensively. So it is absolutely certain that someone, somewhere, will use reinforcement learning to make models very good at this, once coding agents exist.

Even if you are only interested in using this defensively in practice, you can’t really understand it without knowing how both sides work. So if you want to defend yourself, you need to train for it (or pay for someone who has).