There is no truth for RLHF or RLVR. You can't predict against something if you can't check against the truth.
It's not pedantry. The objective function changes. The optimization changes. THese are real things when training a model, not hand wavy philosophical ideas.