I wonder what legal gymnastics are needed for "I can scrape anything off the web ignoring copyright and build a product from this, but you can’t even display what’s on my webpage elsewhere”.
Perhaps that is it in fact. The act of protecting it from scraping means you object. 99% of the blogged contents etc. Big AI helped themselves to was just… there. Public. Not free from copyright but still not paywalled.
Transformation.
Taking something someone else made and showing it as-is, bypassing their own restrictions: No no.
Taking something someone else made, modify it or use parts of it in some bigger thing or completely change it: Fine, if you have money and/or run a company
except a bunch of paywalled stuff did end up in training corpora
Precedent is pretty clear: competitive uses bad, transformative uses good. Xcancel scrapes and then competes directly with X, whereas LLM labs scrape the internet to make an agentic intelligent bot, a transformative use of the scraped content.