logoalt Hacker News

sphyesterday at 2:16 PM5 repliesview on HN

What protection do LLM search engines have against training off content generated by other LLMs?

Will we get to a point where AI-generated sites make up a majority of the internet, and LLMs are training upon their own regurgitations, with exponential amplification of all their lies and flaws?

Or will the pre-2022 corpus human knowledge be considered the low-background steel standard, and anything after that less and less reliable unless certified that it has been created by a human mind and untainted by hallucinations?


Replies

creaturemachineyesterday at 2:25 PM

I have a feeling we're already there.

show 1 reply
gdulliyesterday at 2:31 PM

> What protection do LLM search engines have against training off content generated by other LLMs?

You're talking about a scenario that won't blow itself up in the next few quarters, so it's of no interest to them.

coldpieyesterday at 2:39 PM

I've mostly stopped using the Internet to learn new things and have gone back to books from the library. The majority of technical books at the library were published pre-2020s and hopefully, publishing slop physically won't be profitable enough to flood that market, too. Now that the Internet has largely been destroyed by slop manufacturers, whether or not the words are(/were) worth putting on paper becomes a useful discriminator.

kjs3yesterday at 6:27 PM

Will we get to a point where AI-generated sites make up a majority of the internet

I dunno if they'll be the majority (I suspect we're alredy close to 'yes, and it's already happened'), but I feel very, very confident that they will be the majority, if not the totality, of sites that the vast majority of people see.

NegativeLatencyyesterday at 2:22 PM

They’ll train on prompts and anything else you send in. Many LLM responses are sorta finger printable: I assume this is intentional