I'm stunned that people are taking this accusation as a fact.
OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around on some kitchen table somewhere and 2. I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business.
There are things that Buckmaster alleged and things that he speculated. The entire training data thing is speculation. If this is pissing you off, then you ought to evaluate how you ingest information.
This is how internet discourse works on Reddit/Twitter/HN and the rest. Someone said something which confirms your biases so it’ll now be treated as a fact and repeated endlessly in the echo chamber.
Whenever I see comments defending AI companies, I look at the account's creation date, and interestingly almost all of them were created post 2024.
He asked whether they used their chats as training data and received no response. Any speculation here seems quite appropriate?
I don’t think you understand how brazen big tech companies are in practice.
They stole it.
He didn't even make that accusation!
> I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
The shocking/interesting thing would be if it was trained on the sessions. I think it's very implausible that they gave the model access to someone else's sessions as input. That would be a huge privacy violation and would probably blow up a large proportion of their enterprise business.
Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.