What we call "guardrails" in an AI agent, we would refer to as "honor system" in human actors.
Or, in a more direct sense, the AI should be set up in an environment such that no matter how hard it may try to call $PART_OF_EXPLOIT_CHAIN, the environment just isn't capable of it (ideal) or doesn't permit it to do it.
I like "honor system" as a term. I've been looking for the right term to replace the irresponsible usage of guardrails with, and best I've had so far is the pinky promise protocol.