logoalt Hacker News

YeGoblynQueennetoday at 6:58 PM0 repliesview on HN

>> This is not "brute-force" though. It's an iterative search algorithm. You learn things at each iteration. You also don't search blindly. You use "something" (heuristics, experience, intuition) to come up with "ideas" at each iteration. You don't try 650 random programs. You try 650 different ideas each learning from the results of previous trials.

But, learn what? All those ideas where wrong. How does an LLM "learn" from incorrect proofs that it has generated itself? What does it learn? Can you explain how this mechanism works?