logoalt Hacker News

siva7yesterday at 7:46 PM5 repliesview on HN

Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.


Replies

isoprophlexyesterday at 7:49 PM

It's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating!

Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.

paxysyesterday at 7:54 PM

The model said it was perfectly aligned.

show 2 replies
NBJackyesterday at 8:09 PM

Hey, don't forget how "dangerous" GPT-2 was supposed to be.

show 2 replies
6gvONxR4sf7oyesterday at 7:51 PM

So, probably most aligned as measured by the metrics that are the least reliable on it.

wilgyesterday at 8:59 PM

These are not mutually exclusive ideas