logoalt Hacker News

disgruntledphd2today at 7:18 AM3 repliesview on HN

I'm not really convinced that there's much secret sauce here, all the methods and data are public, the only real difference is how much compute it takes.


Replies

didroetoday at 10:14 AM

All the methods and data are not public. We don't know what unpublished methods they're using. You can get most of the pre-training data publicly but they've probably spent a ton of money curating it and are now doing things like buying rare books. The RL training data is all (/mostly) proprietary though, and that's the real secret sauce part.

show 1 reply
Amekedltoday at 8:30 AM

everybody do be cooking with water. Chinese Labs provided pretty good, primarily cost-reducing techniques, like the sparse attention patterns recently. I'd bet OpenAI and Anthropic use their variants of those too, so they can get greater margin on their tokens - not something they'd really want to / need to self-report.

cubefoxtoday at 9:21 AM

This is obviously false. There is "secret sauce" because in fact not all the methods and data are public.

show 1 reply