Me and a Cofounder are building something along these lines on top of K8s. The wisdom that "if give the LLM clear boundaries and you get better results" definitely holds IME, even with frontier models like fable. It's impossible to specify all the guardrails you need to keep models from violating priors without writing the code yourself, so the only move is to remove their ability to do so, or to even recognize that the option exists.