There is some aspect of baking the rules into the model, that's what I refer to as RLHF above (which I use here as more a catch-all term for a variety of post-training activities), but there are also external to the model classifiers that may run on the prompt input or on proposed tool calls to limit the way in which the model is used.