The more important question - how would you actually get LLM agents to follow the instructions in your ever-growing CAPA reliably?
It’s all very well having a list of actions to avoid but that doesn’t help if your agents won’t reliably follow it.
Why couldn't you express all those as test cases rather than instructions?
In test cases i can do anything, a test framework is just a way of discovering and then scheduling functions to run. I can emit useful instructions to the agent from the failed test case: "After walking the AST of all use of state machine X, a branch was found at Y which reused stale state. Ensure stale references are dropped..."
I can force the agent to pass the test suite before it considers itself done. I can reject edits of such test cases to partially mitigate reward hacking. etc etc
This is a question of context management and I suppose some would classify this as "harness engineering" as the trend of the moment.
One approach, for example, might be to have the a standalone code reviewer agent that is solely responsible for interfacing with the CAPA system (e.g. via a tool, via MCP) and acts as a back stop. When it finds a new type of CAPA, it stores it (and the backend indexes it with enough metadata to support broad types of retrieval). When it reviews a piece of code, it finds past CAPAs. By file locality. By business domain in the application. By keywords.
Same tool and repository available to both building agents and review agents, but use the review agent as a dedicated back stop as part of the verification process.
I think the only way for now it through distillation of your own models, which can get expensive fast.
The same way you do for humans, regular training and audits.
You stop relying on the agents following instructions exactly.
You need two pieces:
a) prompts, that tell the agents what to do and how to do it (and ideally, the why, where, etc, the full picture) - that's the positive half, that drives behavior the way you want it.
b) deterministic tooling that prevents negative outcomes, like linters, compilers, static analysis, fuzzing, testing, the more the better. This side should either be firewalled off from the AI or very carefully watched so that it doesn't drift.
The part that you put in the deterministic side is the "never do x" stuff - I have lint for long comments (which AI hits every single time it commits), all my dev scripts are in typescript, precommit hooks, massive CI, and I lint even for things like redirecting error to standard out, tiny stuff, and also e.g. static migration analysis so the AI never ships an exclusive full table lock in a migration, for example.