There's a difference between automated industrial scale scraping and well-behaved agents acting on the behalf of individuals. Right now there isn't a robust, standard way to distinguish between them, so sites just block known datacenter IPs and throw out the baby with the bathwater.
That's a main advantage of running your own local claw setup using your residential connection -- difficult/impossible to block.
That said, eventually a site blocking all agents would be like blocking all search engines, something that hurts more than it helps as agentic interactions become "the norm". WebMCP or similar support will likely be a basic expectation at some point.
> Right now there isn't a robust, standard way to distinguish between them
There is, it’s called an API
> That's a main advantage of running your own local claw setup using your residential connection -- difficult/impossible to block.
It's worse than that, they will block all agents and specifically make exceptions for maybe 2 search engines per country. Thereby strengthening existing players and weakening all incumbants.
Cloudflare for instance does allow indexing by Google and Bing I think by default, but does challenge other bots. Double check me I am speaking from memory.