logoalt Hacker News

eloisiustoday at 10:59 AM1 replyview on HN

That (theoretically) solves training, but it doesn’t change the fact that even smart models can’t extract useful information from a dead internet, so you’ll always be stuck with a stale training cutoff. This is already a problem I run into a lot. I search something first. Top results are slop sites, so I switch to a chatbot. Its answers look suspiciously similar to the slop sites I just noped out of. Check the sources. It’s them.


Replies

backlava12today at 1:58 PM

And the training of future models will have to contend not only with slop, but also huge amounts of content specifically designed to "taint" future training data. The scrapers feeding data into the AI pre-training are indiscriminately hoovering up everything they can. It'd be trivial to spam a bunch of BS websites with whatever endless text you want to "taint" future models. Post tons of examples of insecure code, publish package.json files pointing to some malicious library, etc...