logoalt Hacker News

ardel95yesterday at 4:33 PM1 replyview on HN

CSAM, and other harms, are typically detected using a set of specially trained, faster and cheaper models (and out of band matching techniques) that run before and after the main model.

Any mention in the system prompt is mostly defense in depth, and to make refusals more graceful.


Replies

whstlyesterday at 7:06 PM

Also, the system prompt, or even something reinforced on every message, is nowhere near as strong as its internal training or as an external safeguard.

If the prompt were the only protection, it would be extremely easy to produce illegal content after a long session.