logoalt Hacker News

8noteyesterday at 8:20 PM2 repliesview on HN

alignment isnt particularly required

we are passing in training data that says to do those felonies. we dont have to. we could also have the thing predict whether what its about to do is illegal or not before doing it.

theyre choosing to build felony harnesses. the model just outputs tokens, not felonies


Replies

cowanon77yesterday at 8:29 PM

> we are passing in training data that says to do those felonies.

Partially, but also I don't think current AIs really have any judgement of right and wrong, they just see chains of reasoning between ideas. This is the deeper issue, there is no way to sanitize the data or training to fix it. Current AIs are fundamentally unsafe, and only become more unsafe as they become more powerful.

estearumyesterday at 8:29 PM

Assuming "adherence to arbitrary, implicit, and context-dependent rulesets" is the default behavior of uhhhh... anything at all... is a truly ridiculous assumption.