There is a command in oh-my-pi called "/omfg <problem>". You explain what is wrong with agent's response, and it writes a hook to make sure that the problem doesn't happen again. It then re-runs your previous prompt to make sure that hook is triggered, and if not, it rewrites the hook to make your previous prompt trigger the hook. Then each next agent's response is checked by the hook, and if it is triggered, the agent receives feedback on what's wrong and what must be done differently.
How does it generally work? Does it run a smaller model on a small set of tool calls/previous thinking block?
I've written a lot of custom hooks this way. It's an amazing feature.
Sounds like something that can grow unwieldy.