logoalt Hacker News

umviyesterday at 11:30 PM7 repliesview on HN

I feel like collecting, curating, and protecting high quality corpuses of "truth" is going to become increasingly important for high quality AI.

There will come a day (and probably soon) when "training on the public internet" (Reddit, etc) will taint your model with metric tons of corporate contamination, political poison, and other adversarial content intentionally crafted to bias AIs for various reasons (corporate gain, geopolitical information warfare, etc). Basically the AI-equivalent of SEO.


Replies

jessetempyesterday at 11:38 PM

All of that already existed for the purpose of biasing people and now it biases ai for free. A company would have to make an effort to remove or change the bias

marcd35today at 7:58 AM

Reddit has already begun the effort to start poising the well - https://www.reddit.com/r/poisonai/

show 1 reply
hadlockyesterday at 11:36 PM

I think most everyone already has a curated training library; Web scraping exists but I don't think anyone is still using it as a primary information vector

show 1 reply
altairprimetoday at 7:50 AM

If I ever curate again it will certainly not be for the public. That led to PageRank which kickstarted this whole dystopian nightmare that Google has been planning since as early as 2003. No thank you.

satvikpendemtoday at 1:06 AM

This already exists, there are archives of Reddit or other sites, and Anna's Archive for papers and books.

asawfofortoday at 12:44 AM

Isn’t this what the paper-bound encyclopedia companies do, albeit shallowly

latexrtoday at 7:36 AM

> There will come a day (and probably soon)

That day has already arrived, it is already happening.