logoalt Hacker News

kbwal7 • yesterday at 8:10 PM • 0 replies • view on HN

Note that this sort of distillation is NOT for pre-training data (which is tens of trillions of tokens). I think the allegations against Chinese companies by Anthropic is more so that they distill SFT data (which is good for post-training, but you still need a strong base model)