Crawling the internet and dumping it to disk is not "stealing".
Then distilling models and deobfuscating reasoning traces isn't.
Mass downloading copyrighted works is. Which they did. Aaron got threatened with 20 years, they got pentagon contracts.
Is everything licensed in the same way? Are there any copyrighted works available to be had through crawling?
It's not stealing but arguing that it's not infringement because its on the internet is pretty obviously nonsense.
If they were only copying, for example, New York Times articles and many publishers to a disk, I don't think NYT and the publishers would have sued OpenAI. But OpenAI isn't just copying things to disk. NYT reported ChatGPT (before Dec 2023, [0]) was returning near verbatim sections of NYT articles.
Is this stealing? Is it depriving NYT or publishers/writers from money via lost sales/subs? I don't know, but it certainly could be.
[0] https://www.nytimes.com/2023/12/27/business/media/new-york-t...