logoalt Hacker News

alightsoultoday at 7:12 AM2 repliesview on HN

GitHub dumps are about 115 terabytes. The common crawl is in the petabyte range uncompressed for every year. Apparently there are dumps of Reddit too in spite of their efforts to ban bots and it's not solely due to the use of residential proxies. For a 1:20 parameter to token ratio, you can still train up to 10 trillion parameters so 10T parameters times 20 is about 200 trillion tokens. Then each token is 4 bytes so 200 times 4 is about 800 terabytes, which is not inconceivable, the common crawl alone has more data than that. So does the internet archive if you donate to them, Anna's archive is 2 petabytes including images, etc etc not all of it is text, but training on multimodal data increases model intelligence by virtue of being multimodal


Replies

alightsoultoday at 3:08 PM

also reddit has eliminated their api entirely, but dumps of it can still be made. every website can be seen as its DOM with html, css, javascript, which can be seen as source code especially if you only look at its javascript, and its dom with css, html, javascript or only javascript can be added to a source code dump together with github and can be duplicated as plain text with no html markup, no css, no javascript, as an information source. if you pay youtube, instagram, tiktok, bilibili to crawl their data, you can probably get data into the exabyte range.

miohtamatoday at 8:38 AM

Maybe Reddit dumps explain why Opus 5 is talking like a retarded.