logoalt Hacker News

stymaaryesterday at 8:07 PM1 replyview on HN

> Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data.

What was the difference between what deepseek did for R1 and what OpenAI did for o1?


Replies

npntoday at 3:24 AM

openai did human crafted chain of thought dataset training. deepseek didn't have the resources so they attempted RL. doing RL correctly is hard because of the risk of model collapsing.