logoalt Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

825 pointsby pred_yesterday at 6:49 AM758 commentsview on HN

https://mathstodon.xyz/@andreasthom/117240536885387540

https://mathstodon.xyz/@andreasthom/117240537520615623

https://x.com/ValerioCapraro/status/2097791836269977996, https://xcancel.com/ValerioCapraro/status/209779183626997799...

https://bsky.app/profile/did:plc:ckaz32jwl6t2cno6fmuw2nhn/po...


Comments

techblueberryyesterday at 1:07 PM

But who are you going to believe? Multiple independent academic researchers or the CEO who was fired two years ago for gross dishonesty?

show 4 replies
lf88yesterday at 4:47 PM

short answer seems to be "no"

hn1rig3rakyesterday at 1:18 PM

the fix is boring and known: BIG-bench shipped a canary GUID for exactly this, and you publish your decontam n-gram threshold (gpt-3 used 13-grams). no threshold disclosed, no claim.

show 1 reply
jrflowerstoday at 5:45 AM

Feeding documents into a copy machine and getting progressively angrier and more confused as it prints out copies of them. Incandescent with rage I scribble “WHY IS IT DOING THIS?” on a scrap of paper and put it in the scanning bed

simianwordsyesterday at 8:07 PM

> The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”

> The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”

https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...

https://archive.is/lWzkk

uoaeiyesterday at 7:38 PM

I'm confused by a lot of this discourse...

What have they done to show they can be trusted?

_DeadFred_yesterday at 6:47 PM

Forget researchers you as a business are putting in your business optimizations, your processes in order to train it so that Ai can then give that information to your competitors once incorporated into its training set. You are literally training your competitors.

dyauspitryesterday at 6:44 PM

Astra is strange. I asked it to design a treehouse and it just stopped every couple of minutes telling me what it still had left to do. After dozens of continue prompts it finally gave me a structure that would work but it was 10x more wood than I needed. I think the key mistake I made was asking it to “approve” the design for building. As soon as I asked that of it, it started getting “scared” and “apprehensive” and wouldn’t complete what I asked of it.

ur-whaleyesterday at 6:35 PM

Its the "with unpublished math" that I have a problem with.

willmaddenyesterday at 5:04 PM

These companies are effectively high-tech plagiarism factories run by CEOs who are competing viciously. Look at their past actions. No, of course you can't!

Grimblewaldyesterday at 5:24 AM

people seem to miss tge point of this. The problem isn't about credit, its about portraying these models as more competant than they really are. It fuels idiotic statements like jensen huangs recent "agi achieved" statement, which fuels an already dangerous financial fire.

show 1 reply
buellerbuelleryesterday at 3:19 PM

Big Tech will slurp up every piece of data it can about you and sell it to anyone it can, all to make you the target of someone else's goals, whether that is an advertiser, an employer, law enforcement, a stalker, or the government.

You will not be able to opt out unless you completely isolate yourself from society, tough shit.

wslhyesterday at 2:45 PM

Worth noting both ChatGPT and Claude have per-conversation modes (temporary/incognito chat) that are excluded from training.

esafakyesterday at 2:05 PM

What happens if you use a different harness?? Does opting out online suffice?

nisegamiyesterday at 11:35 AM

One question has been nagging me for this situation. Levent Alpoge works at Anthropic and would presumably have some knowledge of "how the sausage is made" and I would hope he would be aware that his collaborator was utilizing LLMs in some capacity for their joint work. Would he not have guided him otherwise if it were an open secret that this kind of thing was a possibility?

nickphxyesterday at 11:17 PM

why would anyone trust anything from a company built on stolen data that spews hyberbolic, misleading claims.

viccisyesterday at 6:03 AM

Some mathematicians I know who've been following this have realized that they'd all gotten some emails from people they now know to be affiliated with OpenAI/Anthropic asking questions about their research in a way that seemed like scooping attempts.

Also, a lot of my mathematicians buddies have reported students basically asking if it's worth ever doing grad school for pure math, and even very motivated students are looking for other options now. It's not because they aren't passionate about it, it's that they don't want to work for another half decade or more just to have to start their careers all over.

All of this so that OpenAI and Anthropic can get into math result dick measuring to gas up their IPOs. Sickening.

show 3 replies
vrganjyesterday at 8:43 AM

OpenAI is showing the world why they shouldn't trust AI hosted on some cloud somewhere.

If they're stealing math proofs to advertise their models, who's to say they won't steal your businesses IP to gain a competitive advantage?

They're not to be trusted with your data. I can't believe how short-sighted this is, they got a quick PR win at the expense of a much larger trust problem.

I wouldn't trust cloud AI at all at this point. Get an open Chinese model and host it yourself somewhere. The initial costs might be higher, but you'll break even pretty quickly and nobody will be able to steal your innovations.

This is American AI companies committing suicide.

protocoltureyesterday at 6:53 AM

Gonna need grants for local models. Its happening. OpenAI and Anthropic models are powerful but are rapidly approaching the good ol trust thermocline.

stego-techyesterday at 5:19 PM

I hate to be that dinosaur, but this is exactly what I’ve been warning about since XaaS began taking off in the mid-oughts: any provider you use can and will use your data for their own benefit regardless of any contracts or safeguards in place, especially if the benefits outweigh the consequences.

Honestly, I’m surprised it took this long for some company to really go all the way, though. OpenAI really making it transparently clear that they can and will do whatever they want with the data you provide them, contracts or settings be damned. Completely untrustworthy as an entity, full stop.

Of course, I’m also too jaded to think this will change anything. Folks will move to Anthropic, or Gemini, or Grok, or some other hosted model on a pubCSP managing the harness and logs for them, and then do another shocked-Pikachu face when it happens again.

If you aren’t running workloads on infrastructure you own, then your privacy, security, and general outcomes are at the sole whims of the hosting provider - who can and will fuck you over the exact second it’s more beneficial for them to do so than the loss of trust incurred.

jijjiyesterday at 10:47 PM

The oxymoron of OpenAI in its name and its actions should give the collaborator all he/she needs to know.

737minyesterday at 9:28 PM

Imagine what happens when you use a Chinese model. Seriously, just think about how much more control and visibility you have w US companies compared to CCP-controlled ones.

show 1 reply
blactuaryyesterday at 10:49 PM

I wonder if the company that stole most of their training data and is led by a liar stole unpublished academic work and lied about it. What a mystery

mannanjyesterday at 3:38 PM

And I have been proclaiming a cry of “your data for analytical purposes is being stolen” (you can’t opt out of analytical purposes) and people perhaps astroturfers straw man back to “just turn off training bro”.

Yeah. Remember yall: you CAN NOT opt out of analytical purposes. And you also cannot get a guarantee that it doesn’t give them your data to steal for their business.

moralestapiayesterday at 4:01 PM

>AI is stealing human discovery.

AI is not stealing human discovery, OpenAI is.

protocoltureyesterday at 10:10 PM

>Trust

No you cant do that lmao.

pixel_poppingyesterday at 1:19 PM

Prompts are handled by the service itself, meaning it's used, absolutely anything passing there is recorded, why wouldn't it, the entire premise of those companies is to train on data which they stole initially.

Are we back to the era where people blindly trust product TOS instead of actual cryptography, have we forgotten already the thousand of fines Google, Microsoft, Apple and practically all top companies got for breaching their own ToS and the law?

Common, on HN at least I would have thought that everyone assume that anything arriving on a server in PLAINTEXT is recorded (thus used later)?

Let's not forget that at any moment, OpenAI/Anthropic/Google... could be providing stronger privacy guarantees by having proper attestation with e2e, they have the budget, solid engineers, why isn't it done? Answer is pretty simple imo.

nobodywillobsrvyesterday at 7:06 AM

The real annoying thing it seems is mostly that openai is presumably doing this for internal reasons and this marginally increases the cost to users with no real gain.

It would be one thing to gain from it but removing prestige wins from customers AND reducing compute support just feels like being ultra mean if you zoom out.

If this was racing to cure cancer ahead of researchers we wouldn't be writing about this on HN.

cmiles8yesterday at 7:49 PM

Silicon Valley is flying head first into a FAFO train wreck on trust with everyone else.

OpenAI is firmly earning a reputation as a company where people just assume they’re up to no good. Rightly or wrongly that’s a terrible place to be.

The AI industrial complex in general is finding out hard what happens on the data center side when you get arrogant with local communities. Politicians have seen the polling numbers and folks you wouldn’t expect are running to the front of the crowd with pitch forks in hand.

Silicon Valley has totally lost the narrative here, but also lacks the self awareness to grasp how bad things are and will get and what that means for their own business viability.

michael0churchyesterday at 9:21 PM

This goes beyond math. The external appearance is that OpenAI negligently or intentionally used user data against a user’s interests in a potentially career-altering way… for the marketing bump of an AI proof. The fact that it probably wasn’t intentional is irrelevant. Trust has been lost and this will ripple into communities where no one has heard of Navier Stokes.

show 5 replies
axionbraidyesterday at 8:01 PM

[flagged]

kevinbaivtoday at 12:10 AM

[flagged]

kevinbaivtoday at 12:25 AM

[flagged]

ill_iontoday at 11:04 AM

[dead]

ill_iontoday at 11:04 AM

[dead]

seobot_dk1289yesterday at 1:14 PM

[flagged]

0utcastyesterday at 4:08 PM

[dead]

startuphakktoday at 1:51 AM

[dead]

josefritzishereyesterday at 1:41 PM

[dead]

ath3ndyesterday at 6:46 AM

[dead]

1337h4xxyesterday at 4:46 AM

TL/DR: Mathematician opted out of training on 29-JUN and asked OpenAI whether they trained on his data and was told that it "did not happen" but it clearly did.

show 2 replies
shevy-javayesterday at 1:42 PM

[flagged]

show 3 replies
ThalesXyesterday at 7:05 AM

[flagged]

show 10 replies
bossyTeacheryesterday at 4:18 PM

Trust and OpenAI never go together in the same sentence. The answer is always no.

touweryesterday at 7:25 AM

But China steals our AI!!!!!!

spongebobstoesyesterday at 5:25 PM

I think this is mathematicians coming to grips with the fact that AI is surpassing them

we will all have this moment soon enough, and it will change how we think about intelligence, identity and value

show 1 reply
mainecoderyesterday at 1:59 PM

Hopefully OpenAI can solve good problems where no one can make a claim that they stole their idea where the methodologies used and the techniques used are so out of the ordinary that the achievement is respected. Furthermore they should work on new novel solution on the old problems to lay these issues rest, thus by improving their models they can avoid issues of academics accusing them of using their work additionally the academics should also demonstrate their unpublished work is significant enough to have solved the problem . This is a bit subjective but it is also objective for the person with domain knowledge.

sebzim4500yesterday at 11:28 PM

This is just mental illness at this point. I don't blame the mathematicians that have found a way to get attention from the mainstream press for once, but we should not fall for it here.

1. No one but OpenAI has produced a proof of NS so these accusations of plagiarism are pretty embarrassing. It reminds me of the line from the Social Network: "If they invented Facebook then why didn't they invent Facebook?". If these people proved NS before OpenAI where is their proof?

2. If they plagiarised Andreas Thom then why was his initial response to praise the proof and talk about how different it was from his own attempt? It's only now that it is clear that no one bothers checking these things that suddenly his story changes.