Maybe someone can explain why RL is even needed for post training with Jev? We have supervised labels.
I guess it's due to the calibrated decision part (and that's what LLMs tell me).
But I figure some supervised classification post training would still improve the model.