logoalt Hacker News

KaoruAoiShihotoday at 12:36 PM8 repliesview on HN

Appears to be benchmaxxing

https://x.com/quietnning/status/2080786711861407883


Replies

jchwtoday at 3:24 PM

I have been claiming that I don't think Chinese AI companies are benchmaxxing harder than American AI companies, which has gotten mixed reception: sometimes people agree, sometimes they disagree.

It seems I was wrong. American AI companies might actually be benchmaxxing harder.

show 1 reply
throwa356262today at 12:43 PM

"Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration."

I guess there is no way this can happen without benchmark being part of the training data??

show 3 replies
jnwatsontoday at 1:05 PM

I was just thinking they need to mark each model per benchmark as "model released before the benchmark was released" and "model released after the benchmark was released".

root-parenttoday at 4:24 PM

And being worst than previous model...

"...The traces tell the why: (1) On our most classic Witness-style game, Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration. It already knows this genre. (2) But on our most novel game (unusual mechanic combinations you can't pattern-match), Opus 5 regresses below Opus 4.8. Where rules must actually be discovered through interaction, the new model is worse than the old one..."

Zababatoday at 4:29 PM

>That decomposition (perfect on templates, regressed on novelty) is the signature of “scaffold-then-internalize” training on genre-specific data, not a general gain in interactive abstract reasoning.

They're smuggling a claim that benchmarks like ARC-AGI measure "interactive abstract reasoning" here, which is what is claimed by the people that make these benchmarks, and also not proven.

andrepdtoday at 1:19 PM

I'm shocked, astonished even, that enterprises on which trillions of dollars are being poured would consider cheating on marketing benchmarks.

show 1 reply
codewiththihatoday at 4:03 PM

[dead]