logoalt Hacker News

elar_veroletoday at 12:11 PM3 repliesview on HN

I think this can't work because an LLM needs too much data, and before the internet there probably just wasn't enough to get close to what we have now


Replies

jvanderbottoday at 12:16 PM

Even simpler: Can GPT-2 anticipate and build Gwen/Deepseek? I think the answer is almost trivially "no", so I wonder what changed?

show 2 replies
inigyoutoday at 1:04 PM

Why couldn't an LLM, if it was smart enough, generate and consume its own data?

I know the answer: because it leads to model collapse. But why is that? Wouldn't a smart model not collapse? It's seeming like they keep getting smarter because we keep pouring more of our own knowledge into them, not because they are actually getting smarter. And yes, sometimes a dumb but persistent bruteforcer can make new discoveries.

show 4 replies
hackernudestoday at 12:50 PM

Maybe we can synthesize large amounts of limited information. I thought that new training data is mostly synthetic anyway.