logoalt Hacker News

Arodextoday at 1:05 PM1 replyview on HN

Then OpenAI should acknowledge that they can't prove they solved the problem independently, and credit the external researchers. It cuts both ways: if OpenAI really needs to access user data, even anonymised, to improve its models, they have to waive any pretention to solve "independently" any problem other people worked on with its tools. Otherwise they (OpenAI) have to firewall/cleanroom themselves.


Replies

fc417fc802today at 6:39 PM

So effectively you're implying that any problem directly adjacent to anything that appears in the training data can't be considered independent and needs to be credited?

But at that point you've circled back around to my original objection. That reasoning isn't limited to user data but applies to literally all the training data which at this point (AFAIK) covers the vast majority of everything ever written.

If we accept that position then what do we make of all the other output? Isn't everything it spits out plagiarized? So then is everyone who uses a frontier model to help them in their research effectively laundering plagiarized work? But then the other researchers involved in this controversy were also using the openai model ...