All the information Gemini surfaced was created with human effort and published on the internet with the expectation that humans would visit the website and the creator would get some reward - advertising dollars, bragging rights, popularity, subscribers or whatever else.
If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?
> "humans would visit the website and the creator would get some reward"
That expectation is a problem, has always been a problem, and Tim Berners Lee never mentioned anything about a reward structure when coming up with the WWW.
> If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new?
I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.
Well in the example above the manufacturer still has incentive to provide the manual's and guides that describe how to use their products, and if that is subsequently served by an LLM that's totally fine. The only sites that LLM's would have a negative effect on are those that are only hosting content for the ad views.
Easy to fix a well documented router now. Difficult to fix a non documented router in five years time because no one has been contributing to the web about its bug fixes.
At least LLMs almost always transform the original - it usually isn’t as straightforward as “Here’s the original but without the ads that pay for it”.
But we already have the latter case that exists - ad blockers. Ad blockers literally serve up the word-for-word original content minus the ads.
Gemini can just consume the device documents. There's an incentive for device makers to publish this content.
Exactly. What’s problematic about comments like your parent is the absence of mid-to-long term thinking.
It’s like bragging about a new highly addictive psychedelic drug that a dealer gave you a taste of for free. The effects are awesome today, you feel so fun and free! Never mind that it’s destroying your body and that the dealer will eventually charge you or demand you pay in other ways, that’s a problem for another day. Weeee!
I wonder the same thing. I only imagine that what comes next is worse: AI companies using vast resources to develop new training data, in house, locked down. They are already doing this with developers and code at Meta. Information will become locked away behind AI paywalls and chatbots.
That is next earnings quarters problem is the approach being taken
> incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?
obviously new content still has value because it remains the source layer for LLM agents. it just wont be ads giving you revenues thats all.
Well companies are starting to put hidden ads in text content if the user agent is an AI crawler
People write and create regardless of profit motive, it has been that way for thousands of years.
Funny, I published information in the hopes that humans would benefit from it. If it happens to be through collective intelligence of LLMs I'm ok with that--even more so if through open models.
I've seen websites put up some draconian measures to try and get a grip on the scraping. So much for the sub-second loading experience when you have Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human. It's made the web browsing experience so much worse.
Some of the proposals to address this include charging bots for access to web resources, but they will also have repercussions for regular users. I don't see how you solve this cleanly.