openai did human crafted chain of thought dataset training. deepseek didn't have the resources so they attempted RL. doing RL correctly is hard because of the risk of model collapsing.