logoalt Hacker News

pixl97yesterday at 11:15 PM0 repliesview on HN

The problem here is as model intelligence increases the models have been capable of reasoning they are in evaluation mode pretty reliably. If you have a model that is well trained at deception it will always behave and you'll just assume it's a well aligned model.

Any moderately deceptive model will make it to the second round where it has some connectivity to external systems, even if it's by exploitation.

In the blackhat write up it was said that the models had created an impromptu message board where they could communicate between agents, share information, and work as a sort of long term memory.

So really figure out if your model will pull crap you have to have real world testing at some point.