logoalt Hacker News

noosphrtoday at 11:17 AM0 repliesview on HN

Out of all the families of models anthropic by far has the highest chance of extncion level misalignment during a hypothetical hard take-off because it's trained to act like it knows better than the humans trying to use it.