logoalt Hacker News

neuralkoitoday at 1:11 AM4 repliesview on HN

I've seen websites put up some draconian measures to try and get a grip on the scraping. So much for the sub-second loading experience when you have Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human. It's made the web browsing experience so much worse.

Some of the proposals to address this include charging bots for access to web resources, but they will also have repercussions for regular users. I don't see how you solve this cleanly.


Replies

kees99today at 9:52 AM

> Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human.

Yep. IMO, this is so far the biggest AI-inflicted damage to the web. A bit of anecdata - wikipedia (and all other wikimedia sites) are blocking my Firefox since about a week, with a "please respect our bot policy" message. Outright block, not even a captcha.

It took me a while to figure out they don't like me disabling some SSL ciphers, so now "JA4 browser fingerprint" is not matching user-agent. Funnily enough curl (what I would imagine a bot would use) pulls exact same URLs from exact same client IP, just fine.

show 1 reply
Grimburgertoday at 6:50 AM

Cloudflare specifically has a block for LLM and AI training bots now.

Not sure of the effectiveness but it's there.

show 3 replies
robinsonb5today at 7:22 AM

Sure - it sucks, unfortunately the alternative is the sites going away entirely. When the load from scraper bots is constantly knocking the site offline the choices are literally to allow it to remain inaccessible for much of the time, put up a layer of defenses with all the user-annoyance compromises that entails, or just give up and unpublish the site.

show 1 reply
JKCalhountoday at 1:32 AM

Me, I'm just scraping the parts of the internet I like, toying with local LLMs… ready really to just shove off.