logoalt Hacker News

fnyyesterday at 5:38 PM1 replyview on HN

Why do we hope to use the same model as its own guardrail?

This approach routinely fails with a single stream of consciousness. I can't count the number of times I've had to talk myself out of doing something stupid.

In the same way, a guardrail could inject thoughts like "...but I shouldn't do that..." "...I must remember to respect..." "...these ants deserve compassion."

The guardrail could even go as far as rewriting the thoughts of a model about to go rogue.


Replies

nullbiotoday at 1:30 AM

Presumably because of performance. It'd work well though, I imagine.