On the other hand, the open weights models could crawl and annotate and rl the training data that Anthropic and OAI did in exactly the same way, and take the exact same legal hits. They use distillation because it's cheaper not to do so.
Hmm, why do you think this is true? One reason I'm skeptical of this is because RL envs are often purchased (and are not publicly available), and this might be a sizable component of why models are getting better.
> On the other hand,
What is the other hand here? They could do slow/costly illegal think #1 instead of the fast/cheaper illegal thing #2 that they currently do?