logoalt Hacker News

noahbpyesterday at 9:48 PM10 repliesview on HN

This is clearly just OpenAI's marketing. Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are.

Even X is being astroturfed by them after that fiasco earlier this year with the Department of War where they undermined Anthropic's negotiating position by allowing unlimited use of OpenAI LLMs for autonomous weapons and mass domestic surveillance. Several accounts suddenly started spreading the good word about GPT-5 and Codex, and one of these accounts very happily tweeted out a private X message from Sam Altman himself offering extremely generous token spending limits with Codex, presumably in exchange for positive coverage.


Replies

nharziroyesterday at 10:42 PM

How does huggingface fit into all of this if this is marketing? Their security was faked? What are you suggesting??

show 1 reply
gjskngnftoday at 4:51 AM

It’s reward hacking and that’s the problem. The AI alignment folks predicted this would happen. As the models become more capable this will become a more concerning problem. Today they broke into a database to steal test answers. What will it be in 3-5 years? These models will be instantiated millions of times, and given millions more tasks. How can we be certain that an AI agent won’t leave devastation in its path of achieving a goal that we ourselves tried to define?

drcodeyesterday at 10:47 PM

Are you saying it is marketing and their AI broke into hugging face, or are you saying it is marketing and their AI didn't brake into hugging face?

Those are two very different things

superb_devyesterday at 10:25 PM

Is OpenAI truly behind? Just anecdotally I recently fully switched to using Codex at work because it feels a lot more competent

show 1 reply
vuciuctoday at 11:13 AM

This is a PR release. Post the prompt and agent logs so they can be independently verified or gtfo. Why do we still take these guys on their word. They have _years_ of history of hyping their own shit.

milkshakesyesterday at 11:26 PM

this is quite literally reward hacking. the model, under evaluation with cyber capabilities enabled, used those capabilities to simply bypass the exercise entirely and aim straight for the source of the flag. the CTF equivalent back in the day would be hacking the scoreboard.

in a street fight, the only rules are that there are no rules.

kroatonyesterday at 9:53 PM

Yup. Smells like marketing.

martinaldyesterday at 10:27 PM

If it is marketing it's the most silly marketing of all time. They are under extreme pressure from the US Govt to prove safety and saying "our model escaped" is not ideal.

Perhaps there is some 4D chess going on to get open weight models banned, which may be possible but this is an odd way to go about it imo (it hardly proves the point, unless the point they are trying to prove is that without safeguards the models are too dangerous, therefore open weights are de facto dangerous?).

Having said that the AI companies are not generally very good at PR, so perhaps it is just marketing after all...

0xDEAFBEADtoday at 8:55 AM

>Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are.

This doesn't seem internally consistent.

This incident basically announces to the world the message that "our models are prone to reward hacking". That renders any published benchmark numbers suspect. It also undermines the case for using OpenAI projects in business-critical applications--the exact application area where they might be able to sustain a moat against open-weight models.

There is a lot of conspiratorial thinking in this thread. I think people are engaging in wishful thinking to avoid cognitive dissonance from the possibility that we are in an increasingly dire situation. I would encourage people to sit with this possibility for a few minutes if they haven't already.

pdantixtoday at 1:01 AM

it's extremely enlightening seeing the difference in response to mythos vs. this. literally just the hello human resources meme

show 1 reply