>So, it's still possible they are also running a bunch of poorly-coded scraper bots through residential proxies, but it seems a little non-obvious why they would be doing this.
why "non obvious"? The easiest explanation is gathering training data sets, and is mostly caused by AI companies guarding their pile of essentially stolen IP (given how little they care about copyright) from eachother, without sharing any competition need to get their own and re-crawl to update it too
Well, why run both a well-behaved, easily identifiable bot that probably already gets them all the data they need (they seem to be spending most of their time and money on getting higher quality data than your average internet scrape), and this crap? Like, it's possible the obvious bots are a smokescreen, don't actually work well enough, and they are actually also reliant on data sources that contain 2 million copies of gentoo's bugs database. It's even possible that they are indirectly responsible for it, by buying datasets from shady sources, but again this requires a few jumps that I would like to see justified by evidence.