All the RL data are exactly public. There are huge amount of distilled data freely available, and that amount is more than enough to train a ~10T model.