logoalt Hacker News

TonyTrappyesterday at 7:52 PM3 repliesview on HN

Exactly my observation as well. They devour absolutely everything, no exceptions. No matter how stupid it might be to digest a source code repository via HTTP. They probably don't even recognize what's inside those pages and that there's an easier way to obtain the same result.


Replies

jeremyjhtoday at 3:34 AM

The crawlers are not AI. The crawlers are deterministic. They are collecting data to train AIs.

show 2 replies
diegocgyesterday at 7:59 PM

Which, as the post notes, it's incredibly stupid. So much for artificial "intelligence"

emsigntoday at 12:01 AM

Makes me wonder how much garbage they actually collect across the web. That can't be good for the quality of the LLM.