logoalt Hacker News

lanyard-textiletoday at 2:39 PM4 repliesview on HN

Ah, yes -- A closed source benchmark that Anthropic paid for that Anthropic ranked highest.

0/10


Replies

smallmancontrovtoday at 2:59 PM

Yep, conflict of interest is the elephant in the room and its absence from the "reducing risks from advanced AI" point list is conspicuous:

* Figuring out how to prevent the incentives of frontier AI labs from aligning with anti-social deployment of AI rather than pro-social deployment of AI

It's not like the AI can simply advise them how to fix this because the labs already understood this risk perfectly well before they had an incentive not to. They put in place organizational structures to control it and then promptly smashed the structures once they smelled money. They already failed the integrity check. Even if their AI told them what they didn't want to hear I'm sure they would ignore it. Maybe they already have.

A_D_E_P_Ttoday at 3:51 PM

Opus 5 scored highest, too, which is just lol

Opus 5 is okay at coding. It is unbelievably awful to chat with, though. Very snippy, snarky, and it loves to "push back" even when it's inappropriate to do so. Besides, where "conceptual" things like science and math are concerned, it's starkly inferior to 5.6-Sol. Where writing prose is concerned, Kimi-K3 runs circles around it.

I refuse to believe Opus 5 is at the top of any non-cherrypicked benchmark, unless it has to do with very narrow coding tasks.

show 1 reply
gekoxyztoday at 3:49 PM

Just by seeing the title I knew it was going to be like the meme of Obama giving himself a medal. Why are they even doing this? Are there still people that trust LLM benchmarks made by the LLM companies? It seems to me that the only people still using Claude are the ones that have it for free at work (me). Some of my colleagues even started using their own private OpenAI subscriptions to avoid using Claude, others are using Gemini flash to decipher what Claude is saying...

jrflotoday at 3:50 PM

The conflict of interest is real there, but some benchmarks really ought to be closed source. Otherwise, the second your benchmark is public labs will overfit their new models on it and it will cease to be useful.