logoalt Hacker News

egl2020yesterday at 5:44 PM1 replyview on HN

Any idea how being persistent is trained? I've noticed that telling an LLM that it needs to think some more sometimes produces better results, but the claim here is that "they are very persistent" and "...kept going...".


Replies

naaskingyesterday at 7:45 PM

It's from work like this:

https://arxiv.org/abs/2309.11495

A RL pipeline can reinforce verification behaviour even better than simple prompting.

show 1 reply