logoalt Hacker News

mewse-hnyesterday at 5:39 PM11 repliesview on HN

"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."

What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?


Replies

WarmWashyesterday at 6:10 PM

Everyone knows that they train on the discounted rate plans data. All the labs are upfront about this too.

If you need privacy, then you are going to have to pay full price for those tokens (API). This has been true since day one. Everyone knows it, I guess though this is the first time that it has become "real".

show 5 replies
nullbiotoday at 5:18 AM

I've been saying it for a while now, but no one gives a fuck. Let me repeat it again.

THE BIG LABS CLEAN ROOM YOUR DATA (CREATE SYNTHETIC DATASETS ON IT), EVEN IF YOU OPT OUT, SO THEY CAN BYPASS COPYRIGHT LAWS AND THEIR OWN LOOSELY WORDED TERMS OF SERVICE.

"TOS: We don't train on your data" -> Correct. They train on the synthetic version of your data.

I guess we're just going to ignore this forever though. Who cares about the gaping hole that exists in copyright and contract law now that never existed before LLMs were a thing.

show 1 reply
nradovyesterday at 5:46 PM

Is it spying? I think this usage is disclosed in their terms of service.

show 1 reply
mzstoday at 4:54 AM

This is precisely what I would write after just learning that yes it did.

dash2yesterday at 5:44 PM

If they had agreed to let OpenAI train on their data, it wouldn’t be spying.

show 1 reply
elwellyesterday at 8:04 PM

Isn't this a proof that the usage data is truly "de-identified"? If OpenAI could prove that "their usage" influenced the finding, then it wouldn't be de-identified. (Also, it's a bit disingenuous to trim the "While unlikely," prefix.)

show 2 replies
sinuhe69yesterday at 7:16 PM

More like helped improve our work (the disproof)

vessenesyesterday at 7:37 PM

If those researchers did not opt out then training data might go in. I think it’s a courteous acknowledgement; as was reaching out and examining the direction of proofs themselves. At stake here is a particular mathematician dynamic - ego, prize money, and the sense of proprietary ownership that some might feel working on a problem.

All that was just kicked in the teeth by a group with a lot of compute that was like “bro I heard on twitter that Navier stokes could be solved. Let’s try it.” That’s an existential level of engagement that almost no mathematician in history would like.

kyproyesterday at 11:03 PM

I think OpenAI are correct that it's worth noting, but realistically any relevant usage data they have and used to improve their models would be very insignificant unless they were deliberately using logs from other researchers and training specifically on it (which they seem to deny).

The fact the proofs differ suggests that the models were not directed to be particularly focused on that avenue of research nor trained to converge in that direction.

I get the scepticism, but I feel some of the accusations here are bad faith.

jimbob45yesterday at 7:33 PM

What does it matter? They offered concurrent credit to the other team. I thought I saw sole credit elsewhere in the leaked DMs on Reddit too. This is plainly fair.