logoalt Hacker News

doodlesdevyesterday at 10:43 PM1 replyview on HN

Adaptive reasoning is known to be an extremely hard problem to solve, though. It requires you to predict whether a certain LLM, with a certain effort level, with a certain prompt, will give you the right answer.


Replies

lilytweedtoday at 7:12 AM

This feels like exactly the kind of problem domain that belongs in (and can be solved by) RL?