logoalt Hacker News

neziyesterday at 6:41 PM18 repliesview on HN

I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical.

Now, OpenAI is claiming that the model it used to generate the result was not trained on these collaborative communications with the researcher. This is a technical argument that is impossible to verify as an OpenAI outsider, and probably difficult to verify even for internal OpenAI employees. Provenance is hard to track - you would hope OpenAI has very good tools for this, but a full data trail of all inputs is difficult to trace through.

Another interesting thing to consider is if instead of OpenAI doing this, it was another research mathematician A using an OpenAI model just like the internal group at OpenAI did to publish these results. What if the model A used was trained with unpublished communications with other researchers B who were working on the same problem? Should researcher A technically include B as coauthors? How could they do this when they do not know the communications B had with OpenAI? In this scenario OpenAI, as a middle man, has laundered information from B to A, stripping out attribution. A scooped B without even knowing it!


Replies

jjwisemanyesterday at 6:56 PM

First, OpenAI is not claiming that the model wasn't trained on those sessions. What they've said is “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.” and “We did not use their prompts or proofs to prompt our models or direct our agents.” and “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.”

They also said “Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge … and Tristan Buckmaster….” They say the rumor was that two Millennium Prize problems had been resolved, and that this prompted them to launch "an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems."

It's not obvious to me that's an unethical thing to do, if it happened as they described.

show 4 replies
jameslarsyesterday at 6:50 PM

> Provenance is hard to track - you would hope OpenAI has very good tools for this, but a full data trail of all inputs is difficult to trace through.

What would OpenAIs incentive for this be? They've gotten away with scraping everything and getting it ruled fair use. It seems like willful ignorance is an affirmative defense today. Why would they want to have some sort of audit trail that could prove otherwise?

show 1 reply
raincoleyesterday at 7:32 PM

The irony is that OpenAI got into this trouble only because they tried to play "nice". They told Buckmaster that he could publish the final result as the author as long as he removed Alpöge from the author list. They wanted to give Buckmaster a chance to be the one solved N-S problem.

While this behavior is highly questionable, if OpenAI just published the final result without notifying Buckmaster first and simply cited his previous researches, there would be no ground for anyone to accuse OpenAI for anything. Their self-perceived "generosity" backfired dearly and I'm sure they'll never make the same mistake again. There is probably a policy forbidding any OpenAI employee to contact external researchers like that now.

show 3 replies
enyonetoday at 9:47 AM

I think OP's analogy is bad. The difference of OpenAI when comparing to human collaborator is the possibility to replicate once learned skill. Imagine if any single human collaborator learns a skill it is immediately a skill of any human collaborator.

Agentlientoday at 4:39 AM

I think this move by OpenAI is crazy. At best, if all unconfirmed accusations are unfounded, they still heard a rumour that someone had solved a huge million dollar problem and was about to make a name for themselves. Then, they decided this was a good opportunity to pour millions of dollars into trying to snag the glory while the researchers were busy cleaning up their notes and polishing the announcement.

That still sounds highly unethical.

show 1 reply
tw04today at 3:48 AM

> you would hope OpenAI has very good tools for this, but a full data trail of all inputs is difficult to trace through.

They have a financial incentive not to track any of this, so why would they?

OpenAI’s entire business model is predicated on stealing other people’s work and selling it to the masses.

cjf101yesterday at 6:56 PM

If the model was trained proper to the conversation with the researcher took place, there'd be no question of tainting the results. But if any amount of training on the model took place afterward, then yes, everything is thrown into doubt (a core problem with considering anything "original" from a model because of how >a % of everything ever written has been used a corpus for the training).

ameliusyesterday at 10:55 PM

The problem here is that OAI (and others) pretend or claim that this is uncharted legal territory, where in fact it is very simple. We have a machine that is fed data, and produces new data as a result. If that new data depends (in any way) on the fed data, then from a legal viewpoint it is derived from that data.

Whether they anthropomorphize the operation performed by the machine does not matter. They can anthropomorphize when/if the law is updated to include such terms, but right now they certainly cannot.

show 1 reply
Eji1700yesterday at 10:58 PM

> Provenance is hard to track

Right, which is going to open a lot of doors to a lot of questions.

I don't think there's any legal ramifications on this, just ethical ones about when and how you publish research, but it's yet another point in favor of "if provenance is hard to track, should we be using this for things where it needs to be".

Obviously copyright/trademark is a huge discussion on this, and I could absolutely see this devolving into that as well with how certain findings wind up monetized.

We have a response in this topic from someone claiming to be from OpenAI and linking an article where they, roughly, say "we are sure nothing from the 2 month period made its way into the solution". If that is true, that should mean it is provable, but leads to some more open ended questions like "well what data did it use then?". Is this still okay if someone close to the author did plug data into open AI and it extrapolated it?

Obviously that's probably an unreasonable expectation for these models to track and prove, but it also used to be an unreasonable expectation to scrape every single piece of digital and physical info for consolidated data.

If I opine to a friend on a park bench about a story I'm writing, do they get to pull it from the flock feed, shove it in the model, and then provide it to disney?

Legally, right now, probably. But there's going to need to be a serious look at laws and standards. Or a major shift in what is and isn't discussed in public if literally every breath and move you make can become monetized.

BobbyTables2today at 2:38 AM

The AI not being human doesn’t escape ethical consideration - OpenAI employees are culpable for what they build.

This was academic research. Could just have easily been trade secrets and proprietary data.

mcmcmcyesterday at 7:10 PM

> I think it's a useful analogy to compare OpenAI to a human collaborator.

Frankly I don’t buy this. It’s not a human or a collaborator. It’s a tool. This is like saying it’s not Microsoft’s fault if they extract a bunch of data from people’s Excel sheets because they willingly put it into the program. Anthropomorphizing software is ignorant and foolhardy

show 3 replies
jimmyddddyesterday at 7:25 PM

So, at my company (and most companies I think), we use confidential in-house versions of the AI software. We don't want any confidential information leaking into the public realm. Are these scientists doing that, or are they just using the public version of the software?

show 1 reply
rolandogtoday at 12:57 AM

Then there's the possibility of indirect training via modern spy devices ("smart" IoT devices like LG TV's) feeding the transcribed ambient conversation data for summarization to an agent [0].

[0]: https://youtu.be/6IFVTcM28KA

lmeyerovtoday at 2:37 AM

OpenAI says deidentified data from the private sessions go into training. (Well, explicitly said they will not rule that out.) That changes a lot of the conversation.

timcobbtoday at 3:40 AM

> but a full data trail of all inputs is difficult to trace through.

Great use case for AI agents

gueloyesterday at 10:29 PM

Bad analogy. OpenAI spent millions on compute to get their result. This is more like if a billionaire heard of your promising mathematical lead and then gathered hundreds of top mathematicians to work on it.

jltsirenyesterday at 10:41 PM

I think it's better to ignore OpenAI here, because OpenAI didn't do anything.

Academic research is a professional field in the traditional sense. Individual researchers are ultimately responsible for their actions. If some OpenAI employees violated academic norms while doing academic research, they should be judged by academic standards.

Scooping someone else's result is immoral but not an outright violation of academic norms. But if you are in possession of relevant confidential information, you are expected to steer clear of the topic. It doesn't matter whether you actually used the confidential information to get your results, because outsiders can't know that. The mere fact that there is a plausible suspicion already puts your integrity into question.

Tenured professors occasionally lose their jobs over similar scandals (but usually don't). If OpenAI wants to regain some goodwill, it should do a thorough investigation that may lead to firing the individuals in question. If it doesn't find sufficient evidence of wrongdoing to justify any disciplinary action, it probably doesn't gain any goodwill either (as it often happens with similar investigations at universities).

And if OpenAI wants to be a trustworthy partner, it should transform into a company of boring gray bureaucrats who provide an essential service without competing with their customers.

show 1 reply