Well, those companies have bots that identify themselves, and you can see what they're doing. Google especially have decades of experience of designing scrapers and seem to be able to the job of scraping the whole internet without causing problems in that time.
So, it's still possible they are also running a bunch of poorly-coded scraper bots through residential proxies, but it seems a little non-obvious why they would be doing this.
(It is worth pointing out that most of this traffic seems to be dumb: it's stuff like getting lost in generated link forests of some web apps or repeatly re-querying the same endpoint on a super-high frequency. This isn't exactly going to give a good return on investment for AI training data, especially since AIUI the main race for LLM performance now is in good quality training data)
>So, it's still possible they are also running a bunch of poorly-coded scraper bots through residential proxies, but it seems a little non-obvious why they would be doing this.
why "non obvious"? The easiest explanation is gathering training data sets, and is mostly caused by AI companies guarding their pile of essentially stolen IP (given how little they care about copyright) from eachother, without sharing any competition need to get their own and re-crawl to update it too