Just to add that the Dedekind cut example seems to stand beyond any RL policy used in chess or go, AlphaZero or AlphaProof. In the classical RL there is an state-action space. If the solution requires jumping to a totally difference action space (that must be created) the local policy stalls. Dedekind cut is an example of a out-of-distribution state-space generation.