With the recent Navier-Stokes controversy, I think there's a credible suspicion that all your IP you run through these models will end up in these companies' possession. OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this.
Obviously it's not possible to run a company whose value is predicated on its IP that uploads said IP to a third party which might get access to it.
This could mean every potential serious customer would have no option but to seek alternatives to these online services.
> OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this.
I think this is being misunderstood. Codex has a toggle to allow your prompts to be included in training data. They’re saying they can’t be sure if the person had it on or off while using Codex to discuss the work.
They’re not saying that some prompts are mysteriously jumping into training data.
Also, there is a large market for AI services which don’t retain anything under any circumstances for enterprise customers.
Or these customers could just use AWS Bedrock...but their current CEO is an incompetent MBA unable to publicly articulate their biggest advantage, in the context of the current AI usage my companies.
You have access to all the frontier models, but...your inputs are not shared with the model vendors...neither are used to train the next model.
Why am I even doing the Amazon board job for them!??
Almost every serious customer is already using ZDR where nothing is retained at all, instead of "anonymized" data.
> OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model).
2023:
"The approach also aligned with the company’s broader deployment strategy, to gradually release technologies into the world for people to get used to them. Some executives, including Altman, started to parrot the same line: OpenAI needed to get the “data flywheel” going."
https://www.theatlantic.com/technology/archive/2023/11/sam-a... https://archive.is/NmO5P#selection-979.907-979.1177
I don't think this has been a big secret.
I thought this was commonly accepted to be the case that companies which sell access to LLMs are also storing and training on the inputs?
I don't mean this as rhetoric, I did not think many people (except possibly those operating under government contracts, and 'normies' who don't know about these things) were under the belief that their IP was kept secret when they use these services.