logoalt Hacker News

Garlef • today at 5:06 AM • 5 replies • view on HN

I think they even more so need deterministic feedback:

I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.

https://habit-hooks.com/

I'm using it to foster IOSP (integration operation segregation principle) for example.


Replies

fxtentacle • today at 8:19 AM

Deterministic feedback is precisely how frontier models are trained. It’s called RLVR. You let the agent run on a problem and then calculate a deterministic score of how well it did. Repeat 1000x times and you can “brute force” a good solution. (Which includes all thinking traces and you add it to your training data.) And then a Chinese model can copy your advance for 1000x less compute. Which is why US labs call this not learning, but a distillation “attack”. It’s an attack on the business model.

➕ show 3 replies
jcjmcclean • today at 7:37 AM

This looks pretty cool but I'd want to be able to setup a bunch of my own project specific code smells, so it's not just a few generic rules. Is that the approach you've taken?

Then for example you could write your own hook and convert existing code smell documentation which agents ignore into a format that works with the hook.

➕ show 1 reply
jghn • today at 6:22 AM

The readme reads like AI slop. Why would one believe the tools would prevent AI slop?

➕ show 1 reply
OJFord • today at 10:20 AM

This sounds interesting but (from the homepage too) I don't understand how it's different than having it run any other linter?

➕ show 2 replies
spacebanana7 • today at 7:55 AM

I suspect strategies like this might be even more powerful with less intelligent agents. Could a 40B model outperform a 400B model with good feedback and instructions?