logoalt Hacker News

Otterly99today at 8:47 AM0 repliesview on HN

I really wonder how much of safeguarding with SOTA models is actually just "Don't do that" in a prompt?

I'm building agent to control machines at work and have to rely on small models, which are especially bad at remembering rules like this, so it is very apparent to me that soft rules like this are pretty much useless. Might be different for these huge models, but still it seems very shaky to me.