What actually is "scaling post-training"?
More RLVR. Give it verifiable problems, if it doesn't find a solution move on, if it does, use that as a reward signal.
More RLVR. Give it verifiable problems, if it doesn't find a solution move on, if it does, use that as a reward signal.