logoalt Hacker News

nradovyesterday at 8:36 PM2 repliesview on HN

Just putting a clause in a publication won't prevent it from being used as training data. Information wants to be free.

The frontier LLM vendors do sell enterprise licenses which contractually guarantee that your prompts won't be used for training. (Maybe they'll secretly violate the agreement but in principle it's legally enforceable.) Scholars and universities who care about credit and attribution will either have to purchase those licenses or run their own private open-weight LLM instances.


Replies

oldsecondhandyesterday at 9:35 PM

Even the $20 tier of ChatGPT has privacy settings that forbid using the user's data to be used for training. The question is, whether this setting is respected.

show 2 replies
Analemma_yesterday at 8:46 PM

I don't like this and I wish it weren't true, but I think the period of "information wants to be free" is coming to an end, it was a relic of a bygone era. Increasingly, making your information free means you're the sucker who is doing free labor for AI companies, or worse, you're helping your competitors. Paywalls, login walls, and rate-limits are going up everywhere: there's the GitLab news on the home page right now, and sites like Twitter, Reddit etc. which used to be publicly-readable are now gated (and Xitter is using the legal system to shut down any bypasses).

I hate this but I don't think there's any going back now that LLMs exist.

show 1 reply