logoalt Hacker News

jackcviers3today at 1:26 PM2 repliesview on HN

On the other hand, the open weights models could crawl and annotate and rl the training data that Anthropic and OAI did in exactly the same way, and take the exact same legal hits. They use distillation because it's cheaper not to do so.


Replies

jatinstoday at 1:29 PM

> On the other hand,

What is the other hand here? They could do slow/costly illegal think #1 instead of the fast/cheaper illegal thing #2 that they currently do?

show 1 reply
FinchNova12today at 5:20 PM

Hmm, why do you think this is true? One reason I'm skeptical of this is because RL envs are often purchased (and are not publicly available), and this might be a sizable component of why models are getting better.