I think they even more so need deterministic feedback:
I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.
I'm using it to foster IOSP (integration operation segregation principle) for example.
This looks pretty cool but I'd want to be able to setup a bunch of my own project specific code smells, so it's not just a few generic rules. Is that the approach you've taken?
Then for example you could write your own hook and convert existing code smell documentation which agents ignore into a format that works with the hook.
The readme reads like AI slop. Why would one believe the tools would prevent AI slop?
This sounds interesting but (from the homepage too) I don't understand how it's different than having it run any other linter?
I suspect strategies like this might be even more powerful with less intelligent agents. Could a 40B model outperform a 400B model with good feedback and instructions?
Deterministic feedback is precisely how frontier models are trained. It’s called RLVR. You let the agent run on a problem and then calculate a deterministic score of how well it did. Repeat 1000x times and you can “brute force” a good solution. (Which includes all thinking traces and you add it to your training data.) And then a Chinese model can copy your advance for 1000x less compute. Which is why US labs call this not learning, but a distillation “attack”. It’s an attack on the business model.