Under current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.
Ah, this dupes https://news.ycombinator.com/item?id=49638353
what if it wasn't even model training? what if openAI mathematicians just took the researchers' conversations and used them as prompts/info/guidance/context to keep working on the problems themselves? why is that not being considered?
Do local inference (especially if you have a high RAM Mac), to ensure your chats don’t leave device.
I must say, there is some weird feeling in knowing that great minds are naive enough to believe OpenAI wouldnt use their chats in any way. If you give a company information it will be used, regardless of laws or promises.
There is no prove in a world the AI companies would give to you ensuring that they didnt train or use the chats.
Why would you need to train a model on certain specific near prove chat if you just query it?
Besides that, its hard to believe that its the case for every "company stole my prove".
FEEL likes Open AI is doing publicity stunt with its new researches
This might be a hot-take, but unfortunately here using AI for your paper was already a bad decision at first.
It doesn't take OpenAI's responsibilities away but I guess the right way is to never feed of use any AI around unpublished content, at the known cost to see it spread around.
As one said, OpenAO is like this untrustworthy colleague that knows everything about everyone at work: the less you tell him the better.
Even when you pay you are the product
If mathematician was already using OpenAI for research purpose and making progress due to inputs from OpenAI's responses, then I wouldn't put it beyond OpenAI's reach to generate different relevant prompts to make progress by itself. Afterall, Model can keep at it for whatever timeline and keep pursuing all possible combinations it can think try.
If you have a business account the terms say they will not train on your data, so that seems like the easiest route to avoid such questions for researchers.
I just don't care. These people are supposed to be smart and I'm not really seeing that
It's been said before, and it remains a concern, that if AI reaches a point where it can do/build/launch anything without a huge amount of human labor, the AI companies have no reason to let you or I extract that value.
And, if they're able to snoop on and learn from your human process that gets from initial prompt to functioning product/proof/whatever their labor to produce that thing is even lower. With their much larger budget than most folks and even companies have, they can pick and choose the most valuable things to pursue.
That's not to say I think that OpenAI is going to steal that roguelite strategy game you're working on, but the companies that own the machines that turn electricity into software (and soon, electricity into hardware designs) have an advantage in any field where they're useful. They get earlier access to newer/better models, they have larger token budgets, they don't have the guardrails you and I run up against.
Employers fantasize about replacing all workers with AI without thinking through that if AI can replace all workers, then AI companies can replace all businesses.
300 billion tokens is like.. $5-25 million giving range of OpenAI ouput prices, I"m sure they pay less at cost so, I wonder if more $$$ in wage hours have been spend by humans on the problem. My feeling is yes?
See https://openai.com/index/ten-advances-in-mathematics/ for the announcement this refers to.
I think this is stupid, for three reasons:
1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods.
2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user.
3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)
So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training?
How do people become that trusting?
The phrasing itself is guilt tripping
I want bunch of lawsuits, because the way things are described now produces perverse initiatives like try to discuss every possible idea that comes to mind with llm and if any of it works later claim the llm stole it.
I would like to see chat logs etc and understand how much of a progress was done by human.
In case anyone from X is reading this, please fix your “open in app” nag screen. For several weeks now, clicking it in iOS opens the App Store entry for X rather than the app, even when you have the app installed.
Never ever trust OpenAI, they are evil.
Thing build from stolen data continues stealing data - for some reason, I am not surprised. ;-)
Doesn't OpenAI have an active court order forcing them to log everything? Can they even legally offer private conversations?
Reminder that there are degrees of "trained on conversations". From John Schulman:
> pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper
> use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this
> use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"
source: https://x.com/johnschulman2/status/2097440545853637108
Tbh it won't really matter soon.
All the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.?
The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all.
That doesn't mean they can't be useful, or that their products are trash, etc. It just means that they shouldn't ever be trusted. Buyer beware.
In this domain, an apparent single unique piece of work is often composed of several breakthroughs. For example, when Andrew Wiles proved Fermat's Last Theorem, he had to develop multiple new pieces of mathematical technology to get there.
The claim here seems to be that the human mathematicians, working with AI, developed technology to go A->B->C. By training on those conversations, OpenAI was then able to encourage the model to go A->B->C->D.
In my opinion that situation should be acceptable, if openly disclosed, because it is in the public interest to make progress on these problems and because AI is clearly an amazing tool for making progress. But the human mathematicians are saying that OpenAI is presenting as if the model got from A->D entirely independently, without acknowledging their background contributions.
What’s the limit ?
Will Microsoft Word publish your novel on Amazon behind your back ?
Will VS Code setup a website with your app idea ?
Isn’t that the whole spiel of these things, you run all kind of text and other data through it and it kind of remembers it and learns from it then it spouts it back out like a human would. Makes sense to me that a training run based on conversations that were fed into the system by users is results in the model learning from these so the model will spit the knowledge back out again, just in a way that’s not directly attributable to the original content (which is the most important step as otherwise it would just be plagiarism). I guess that’s why OpenAI can get better and better as well so fast, people work with it and teach it how to do things by giving it feedback and iterating with it, and all that goes back into the training loop. And training data about millennium prize problems is probably quite spars. Wonder if anyone has tried injecting nonsense science into the training data (e.g. work out a fantasy science theory with names and all kinds of stuff) to see if the model will regurgitate it in a couple of months for other users.
You can’t trust OpenAI period.
This smells of extreme 'cope'. Am I really supposed to believe that all of these problems could have been solved, were right about to be solved, etc. But it just happens they are all getting solved now when AI is getting really good at Math...
"Another researcher[/artist/writer/musician/programmer/doctor/director/etc] says OpenAI trained on conversations[/imagery/books/songs/code/classifications/videos/etc], then claimed breakthrou[gh/original art/bestselling books/chart-topping songs/unique applications/medical advice/free special effects/etc]"
Welcome to the party, with the rest of humanity.
Question: Can you trust the cloud?
No.
I can't wait for OpenAI to do this to companies firing people to free up AI budgets
Theft machines be thieving.
I have a better question? Why would you trust OpenAI or any AI company, at all? Or you crazy?
How can OpenAI figure out how to be trustworthy?
Let's see:
1.) The tool they made is only possible by stealing the assets of everyone on the planet that published them in a consumable fashion online or even in written form
2.) They are destroying books they use to train with
3.) They are totally careless about the potential negative impact of the tool on everything
Just with that already, I don't see why they ever merited any of your trust.
I bet they are willing to take everything given to them and assess it for marketable merit and in the future take action on those items they deem viable.
Shocking a company that stole data to build their AI would steal data to improve their AI.
One of the complaints from the mathematician is that OpenAI cannot tell whether his data has been used as training data. Not many people realise this is a direct consequence of the GDPR.
The GDPR protects PII, personally identifiable information, and the definition of PII does not include “mathematics that only this person can think of”. As long as OpenAI strips out PII and removes identifiers linking the conversation to a person, the GDPR is happy. Without the GDPR, OpenAI might have kept the identifiers with the data, and been able to say whether a specific conversation was in the training data.
It would be really useful if the researchers disclose their notes and/or chats (or the key pieces thereof) so people can determine how close their work was to whatever the models produced.
I mean, now that they’ve been scooped, what value is there in keeping them private? On the other hand, publishing them can bolster their case and help gauge how much the models may been “inspired” by their work.
Eh was ever confirmed they were under ZDR or not by them? Don't like to blame alleged victims but lack of a clear claim after these many days is not a good look. Was ai research allowed, under which guardrails, and what was the policy in place? That translarency would be first step.
OpenAI's ethical and reputational own-goal aside, my big takeaway is that it seems that:
if I'm using Codex to develop some new algorithm (in any space), OpenAI appears to be training its model on my code sessions
anyone using that model (OpenAI or a competitor) might be able to receive from the model a solution that is similar or the same as the one I developed, emerging from the training data
Assume you can't. No piece of paper or promise will protect you against these behemoths.
Remember when we wouldn't give our data to competitors?
There are so many naive academics. They still believe an "opt-out" button.
Navier Stokes was solved by an internal model, so good luck proving it wasn't trained on Buckmaster/Lepöge or other chats.
Academics don't get that AI is a dirty tech bro industry that stole IP via torrents and runs after every surveillance contract it can get.
Interesting that this is already off the front page after just 4 hours.
Feeding documents into a copy machine and getting progressively angrier and more confused as it prints out copies of them. Incandescent with rage I scribble “WHY IS IT DOING THIS?” on a scrap of paper and put it in the scanning bed
I think it seems sensible to _assume_ anything the LLM reads (if you aren't inferencing it) has a chance of ending up in some database somewhere. Regardless of whether you trust the other party its a sensible thing to plan around.