logoalt Hacker News

WarmWashyesterday at 8:07 PM5 repliesview on HN

It would be catastrohic for any of the big labs if it came out that they were training on what was sold as private.

I get this cynical conspiratorial energy, it fits the internet well, but I can assure you most people with even mild business sense would be intensely opposed to this idea. Well, except maybe Zuckerburg, but they don't really do enterprise anyway.


Replies

thephybertoday at 12:13 AM

No it wouldn't.

It wasn't "catastrophic" for the largest of the 3 US credit reporting agencies when their entire dataset was breached. The company is 100% IP and the only value they have was completely copied. Their largest value is to verify identities by the things Americans know (KDB) and after that "single factor of identity" was 100% compromised, the company only got bigger and more contracts.

When there are only 4 competitors in the large scale foundation model business and they all throw caution to the wind because they are racing to own the "$30 trillion TAM" they are all going to make critical security, RBAC, and segregation mistakes.

Both ChatGPT and Claude threads marked for sharing have been indexed in Google at large scale. This is incredibly easy to tell Google crawlers via robots.txt not to crawl those URLs, but nobody at either of these uber unicorns could be bothered to add that one pattern to the one file.

And all of the skepticism here is about verifiability. The foundation model companies are liable for potentially more the companies are worth if found to be violating copyrights of content used for training. They aren't going to make it easier for lawsuits against them by detailing their data ingestion into training pipeline.

show 2 replies
r_leeyesterday at 8:20 PM

exactly, I don't understand how HN doesn't understand this

by this logic every business contract in tech is just a bunch of lies and means nothing and the only way to do anything is to have a server sitting next to you, otherwise it's "someone else's computer"

show 1 reply
ajbyesterday at 9:29 PM

It would be catastrophic if they violated confidentiality blatantly, but that doesn't exclude learning of any description. For example, an ordinary human being can't fork a subagent for a particular client and wipe it afterwards. Humans can't stop themselves learning, so confidentiality can't ban all learning.

Instead, confidentiality includes not literally copying material, not using trade secrets or inventions, and not using knowledge of business dealings for your own purposes.

So, while I fully expect that the big labs don't train on private material to the extent that they do those things in a blatant way, it would not be surprising if they pushed the boundaries. Humans push the boundaries all the time.

Up until now, machines did not have judgement, so if you set up a machine in such a way that you hadn't ensured it couldn't violate contract, you were culpable. But now that they have some kind of judgement, maybe it's enough to avoid liability to tell it to obey the contract, even if you give it incentives not to. After all, that's how it works with human employees, isn't it?

Perhaps now we have machines that understand language, someone somewhere is working on getting them to understand "a nod and a wink" as well.

realusernameyesterday at 9:57 PM

It's going to be the same as PRISM, people will be outraged and business will go back as usual.

(And they totally won't do it again they swear, the contract says so)

20kyesterday at 9:55 PM

I mean the models have literally trained on:

1. Child porn

2. Stolen music

3. Private github repos, before that was 'stopped'

4. Illegally pirated books

Them training on company prompts against the terms of service would be one of the least bad things that these companies have trained AI models on

Why do you think a company - willing to break the law for child porn - won't break the law when it comes to your personal data?

show 1 reply