logoalt Hacker News

simonwtoday at 12:45 AM5 repliesview on HN

The big news here is that Googlebot will be blocked from September 15th onwards by one the "block training" policies, because Google use the same crawler infrastructure for their search index AND for training Gemini:

> Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line with our call for transparency for website owners. Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service).


Replies

dannywtoday at 3:49 AM

Good. Google's approach here is manifestly predator, unfair, and IMO illegal. They deserve to be in court for this behaviour, and mandating owners give consent for AI training or drop out of Google; which is just a non-starter because they're a search monopoly.

That's exactly what antitrust laws are supposed to do, and I hope at least EU regulators take action. Every single Googlebot crawl in your access logs is a trace for damages.

show 3 replies
jofzartoday at 1:01 AM

We had googlebot blast a random customer system and almost cause an outage, this is when I first learnt that google will use it for AI training also. It's honestly kind of frustrating also because you then search on it and theres (was) nothing on how you are meant to "correctly" tell google to fuck off, and not use it like that.

show 2 replies
motbus3today at 9:50 AM

Don't that feel like a threat to businesses who dare to avoid their content being stolen?

miohtamatoday at 5:30 AM

People will use something for search and something needs to index pages, either for LLM or old school search engine.

inigyoutoday at 1:17 AM

[flagged]

show 2 replies