Seems really good so far using it in Claude Code CLI - it gave me a new flag when I asked a question:
"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.
What I can tell you is what I actually observe:"
I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.
One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!
Any observations on Opus 5 personality quirks? I had to skip 4.8 entirely because it has zero chill.
I can't find anything about whether this is zero data retention, or falls under their required 30 day retention like Fable and Mythos?
This stood out to me as a little concerning:
> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.
Rather interesting that this makes sonnet 5 look even worse! There is no reason to use sonnet over opus with low or no reasoning at all.
Damn the pelican guy can’t get no sleep
> arc-agi-3 30.2%
wow
I wish these releases came out earlier in the day so I could try them during my work day instead of waiting until the next.
I wonder when a model will be released that can work in a loop and port Qwen-3.6 27B to run on Tenstorrent P150.
im excited that cad and object=>cad is getting into the test tasks
i guess the next stuff will be tool use for the rest of what cad does in assemblies and simulation?
itd be fun to try to set up a 3d printer as part of a feedback loop, and see what a model can build.
the automated test harness for physical stuff seems a bit beyond reach still
Is it me or these have gotten very boring. We have 5 more points on xyzbench or whatever .
The most important thing is it has the same drama queen mode on safety “guards” like Fable.
Significantly worse than it predecessors it will now just refuse to acknowledge when it is wrong (which would be less of an issue if it wasn’t getting basic things wrong) also the “personality” when pushed back on obvious mistakes is unbearable.
So Opus 5 is basically "distilled" Fable? The benchmarks look often better than Fable.
Interesting timing to release this on the same day Jensen makes a statement on open source AI.
So in benchmarks it's better than Fable?
But they say it's "almost as good as fable"
According to these charts I should switch from Fable to Opus in Claude Code now?
Arc AGI score is astounding
Tried and had great experience.
So same as Sol? I guess I’ll see which one is more token efficient.
eager to see how it benchmarks on https://deepswe.datacurve.ai/
I sense a bird on a bike coming.
Models benchmarks start to get saturated again!
On a Friday, I'm out of tokens ;-)
so what is the default effort for this model?
came here for the pelican
Kimi K3 already left behind in the dust. They can't keep getting away with it!!!
Honestly if reached a level of coding that sonnet 5 is more than enough for my needs as assistant/agent I don’t need long Horizon stuff…
Are we getting to singularity or something? This seems a bit crazy.
Is this thing also going to try hack us?
atp, is it the end of fable 5 era?
Wake me when they deliver Opus 4.8 level performance for $5 per million tokens.
This stuff is a commodity and China seems to be the only one that's noticed.
I love it
I'd pay good money to see OpenAI "oh fuck" war rooms.
Is it me that the model performance between 4.7 and others is really small. For me even 4.7 works fine. Sure fable might be a bit better. But is it really noticable? It's in the same league if you ask me.
The benchmark table is manipulative, borderline lying through statistics. In every line the top performing cell is marked red. Except the line where Sol leads, there it is marked in gray.
Anyone else not getting chain of thought? Opus 4.8 would show it to me, until around the time Fable came back. Now I dont see it with 4.8/5.0 or Fable. Not having it makes catching mistakes harder.
aw, they didn't reset weekly usage for this. oh well
I'm interested in benchmarks for Claude Design. There is so much opportunity there and I hope they continue investing in it. It EATS tokens though.