Adaptive reasoning is known to be an extremely hard problem to solve, though. It requires you to predict whether a certain LLM, with a certain effort level, with a certain prompt, will give you the right answer.
This feels like exactly the kind of problem domain that belongs in (and can be solved by) RL?
This feels like exactly the kind of problem domain that belongs in (and can be solved by) RL?