logoalt Hacker News

Claude Opus 5

1729 pointsby alvisyesterday at 4:57 PM1252 commentsview on HN

https://www.anthropic.com/claude-opus-5-system-card


Comments

tyreyesterday at 5:32 PM

I'm interested in benchmarks for Claude Design. There is so much opportunity there and I hope they continue investing in it. It EATS tokens though.

show 1 reply
bocyesterday at 5:43 PM

Seems really good so far using it in Claude Code CLI - it gave me a new flag when I asked a question:

"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.

What I can tell you is what I actually observe:"

I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.

One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!

show 2 replies
doginasuityesterday at 11:21 PM

Any observations on Opus 5 personality quirks? I had to skip 4.8 entirely because it has zero chill.

arrowleafyesterday at 5:27 PM

I can't find anything about whether this is zero data retention, or falls under their required 30 day retention like Fable and Mythos?

bovermyeryesterday at 5:48 PM

This stood out to me as a little concerning:

> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.

show 2 replies
himata4113yesterday at 5:06 PM

Rather interesting that this makes sonnet 5 look even worse! There is no reason to use sonnet over opus with low or no reasoning at all.

show 2 replies
MasterScratyesterday at 9:17 PM

Damn the pelican guy can’t get no sleep

yusufozkanyesterday at 5:02 PM

> arc-agi-3 30.2%

wow

shockembopperyesterday at 5:56 PM

I wish these releases came out earlier in the day so I could try them during my work day instead of waiting until the next.

arjieyesterday at 6:48 PM

I wonder when a model will be released that can work in a loop and port Qwen-3.6 27B to run on Tenstorrent P150.

whatever1yesterday at 5:22 PM

Where does this leave Fable? I am confused.

show 2 replies
8noteyesterday at 5:42 PM

im excited that cad and object=>cad is getting into the test tasks

i guess the next stuff will be tool use for the rest of what cad does in assemblies and simulation?

itd be fun to try to set up a 3d printer as part of a feedback loop, and see what a model can build.

the automated test harness for physical stuff seems a bit beyond reach still

born-jreyesterday at 7:07 PM

Is it me or these have gotten very boring. We have 5 more points on xyzbench or whatever .

show 1 reply
theplumberyesterday at 6:51 PM

The most important thing is it has the same drama queen mode on safety “guards” like Fable.

b-sidetoday at 12:00 AM

Significantly worse than it predecessors it will now just refuse to acknowledge when it is wrong (which would be less of an issue if it wasn’t getting basic things wrong) also the “personality” when pushed back on obvious mistakes is unbearable.

uramsyesterday at 5:25 PM

So Opus 5 is basically "distilled" Fable? The benchmarks look often better than Fable.

mcastyesterday at 5:04 PM

Interesting timing to release this on the same day Jensen makes a statement on open source AI.

show 1 reply
korabsyesterday at 6:20 PM

So in benchmarks it's better than Fable?

But they say it's "almost as good as fable"

doctobogganyesterday at 5:44 PM

According to these charts I should switch from Fable to Opus in Claude Code now?

inshardyesterday at 5:16 PM

Arc AGI score is astounding

vinishkapoortoday at 5:46 AM

Tried and had great experience.

spstoyanovyesterday at 5:14 PM

So same as Sol? I guess I’ll see which one is more token efficient.

mkurzyesterday at 5:15 PM

Where is the pelican?

show 1 reply
taf2yesterday at 5:10 PM

eager to see how it benchmarks on https://deepswe.datacurve.ai/

hahahaatoday at 12:13 AM

I sense a bird on a bike coming.

Eldodiyesterday at 5:08 PM

Models benchmarks start to get saturated again!

arjyesterday at 6:25 PM

On a Friday, I'm out of tokens ;-)

toephu2yesterday at 6:06 PM

How does it score on DeepSWE?

show 1 reply
firemelttoday at 1:12 AM

so what is the default effort for this model?

nstjtoday at 7:22 AM

came here for the pelican

internet2000yesterday at 6:30 PM

Kimi K3 already left behind in the dust. They can't keep getting away with it!!!

hmontazeriyesterday at 5:30 PM

Honestly if reached a level of coding that sonnet 5 is more than enough for my needs as assistant/agent I don’t need long Horizon stuff…

abc42yesterday at 6:18 PM

Are we getting to singularity or something? This seems a bit crazy.

mihauyesterday at 5:29 PM

30% on ARC-AGI-3

show 1 reply
Uptrendatoday at 12:15 AM

Is this thing also going to try hack us?

arseniitrutyesterday at 7:48 PM

atp, is it the end of fable 5 era?

_pdp_yesterday at 6:55 PM

Wake me when they deliver Opus 4.8 level performance for $5 per million tokens.

show 1 reply
tomlockwoodyesterday at 11:25 PM

This stuff is a commodity and China seems to be the only one that's noticed.

throwaw12yesterday at 5:12 PM

is coding and engineering solved yet?

show 1 reply
shinhyeokyesterday at 10:33 PM

I love it

ismailmajyesterday at 6:39 PM

I'd pay good money to see OpenAI "oh fuck" war rooms.

holodukeyesterday at 8:47 PM

Is it me that the model performance between 4.7 and others is really small. For me even 4.7 works fine. Sure fable might be a bit better. But is it really noticable? It's in the same league if you ask me.

StrauXXyesterday at 5:17 PM

The benchmark table is manipulative, borderline lying through statistics. In every line the top performing cell is marked red. Except the line where Sol leads, there it is marked in gray.

show 1 reply
jakeoghyesterday at 7:54 PM

Anyone else not getting chain of thought? Opus 4.8 would show it to me, until around the time Fable came back. Now I dont see it with 4.8/5.0 or Fable. Not having it makes catching mistakes harder.

petterroeatoday at 4:04 AM

aw, they didn't reset weekly usage for this. oh well

🔗 View 40 more comments