logoalt Hacker News

fxtentacle • today at 8:19 AM • 3 replies • view on HN

Deterministic feedback is precisely how frontier models are trained. It’s called RLVR. You let the agent run on a problem and then calculate a deterministic score of how well it did. Repeat 1000x times and you can “brute force” a good solution. (Which includes all thinking traces and you add it to your training data.) And then a Chinese model can copy your advance for 1000x less compute. Which is why US labs call this not learning, but a distillation “attack”. It’s an attack on the business model.


Replies

balder1991 • today at 2:53 PM

I suppose the same way that the normal programmers and artists would call LLMs an attack on licenses and copyright.

DelightOne • today at 10:42 AM

Are there good open source setups that generate this training data automatically?

Or is that the secret sauce no one wants to share, the edge people see themselves having.

➕ show 1 reply