logoalt Hacker News

Velocifyeryesterday at 3:08 PM4 repliesview on HN

But why don't they just git clone?


Replies

rcxdudeyesterday at 3:29 PM

These crawlers (in contrast to e.g. googlebot and similar better behaved crawlers) are not very smart: they seem to make very little effort to avoiding crawling useless deep trees of generated pages. About the only thing they seem to put a lot of effort into is avoiding blocking.

(I'll note that while these are generally attributed to AI data gathering because of the timing of when they took off, it's not actually obvious who's running these bots. The big players all have crawlers that identify themselves and are reasonably well behaved, but I don't know if anyone has managed to positively attribute these other ones to any particular group)

show 3 replies
lkbmyesterday at 3:09 PM

Because they're crawling a billion webpages, only a tiny fraction of which can be git cloned, and configuring a special case just for that tiny fraction isn't worth the effort (of the crawlers).

DarmokTanagrayesterday at 4:45 PM

vibe coded crawlers run by morally bankrupt trend chasers aren't going to be the most well engineered systems you come across.

acedTrexyesterday at 3:13 PM

Because the crawlers dont care, they are the internets parasites. Their creators care nothing for people or systems downstream of their greed.