logoalt Hacker News

i_idiotyesterday at 9:02 PM2 repliesview on HN

> Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation

The way they describe makes it look like there was an intention to cheat painting it as human/AGI. If you leave a possible path open and it will always find it.


Replies

pixl97today at 5:21 PM

Hence why we talk about alignment and things like reward hacking. There are lots of people that are saying "if we just .... " the model will be aligned, or that we don't need alignment at all. These people are foolish.

paxysyesterday at 9:08 PM

It’s a mistake to apply human morality to this. It isn’t “cheating”, the model is simply solving a problem it has been asked to solve in every way it can.