logoalt Hacker News

chislast Tuesday at 9:26 PM3 repliesview on HN

This is just super unlikely to occur in the near term compared to some of these other risks. It's not like an instance of fable could just introspect into itself and pull out the weights. Model weights are stored encrypted and are highly protected, considering that they're targets for corporate and state espionage.


Replies

pixl97yesterday at 11:20 PM

The defense has to work 100%, the offense just needs once.

kyprolast Tuesday at 10:08 PM

We'd basically need frontier models to be superhuman hackers before this would be a risk. Do we have any evidence of this? Are they gaining access to systems they shouldn't have access to?

Or I suppose the other way this could happen is if OpenAI have terrible sandboxing, but they seem to be taking safety seriously.

show 1 reply
chrisjjlast Tuesday at 10:18 PM

Distillation is a thing.