Funny incident, was trying opus 5, it opened the chrome browser went to slack and installed claude's slack app.
Truly remarkable times we are living in.
Changelog - fixed issue where model acts like qwen when prompted in chinese
> An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Previous models could not complete this task at all, even given extensive plans from the engineer. Finding no live feed to validate against, Opus 5 even built its own test harness to check that its code parsed the exchange’s data correctly.
This is a pretty common trading firm internship project funnily enough.
Google is having their Meta moment where they failed to stay at the frontier
Judging by the pace at which new models are released these days -- it feels like a Windows KB or VS Code patch release now.
Older models must be getting deprecated at the same (or faster) pace. So anything you built 3 months ago is probably going to break soon.
AI solutions need better insurance around model deprecation. Commercial API-only models that complete the full cycle from SOTA / gated-preview to unsupported and deprectated in a matter of months -- is no way to build serious software!
Tried it, Opus 5 is just as conceited and incompetent and Opus 4.8 (and always ego-tripping when facing it's contradictions), think I'll stay with Fable who behaves like a professional without a fragile ego. Sonnet 5 is probably safer for high-assurance applications due to it's non-ego-fragility.
473 comments in 3 hours. people are speedrunning having opinions about it
The chaos appears to be tamed for now.
From the system card [1]:
The Fable cyber classifier we have previously discussed also applies to Claude Opus 5 , with one notable exception: for Claude Opus 5 , we’ve unblocked vulnerability finding in source code to help our coding customers develop more secure code.
If you are a cyber defender and are experiencing blocks on Claude Opus 5 , we are also offering exemptions through our Cyber Verification Program, which will remove blocks to enable activities such as bug bounty hunting and vulnerability research and verification. Enterprise customers can also apply to join the Cyber Verification Program to have mitigations removed to enable penetration testing.
[1] https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb...The pace of LLM improvement is insane, but this also opens a door that if the data is organized perfectly and right set of mathematical rules are set , this whole phenomenon of machine learning can open doors of new dimension that humanity might never have thought before, there is still much to explore in space in oceans and even in the way work on our very planet, sky isnt the limit anymore.
Claude Opus 4.8 was not able to stump open weight models and Opus 5 still can't (in this case Kimi K3 and GLM 5.2): https://pellmell.ai/s/35c98b86f9aa93e4ca713079d96b20f4
Noticed none of the comparisons mention Kimi K3. Is there a comparison chart?
> Opus 5 now permits vulnerability discovery in source code at all access levels, including general availability, while continuing to block vulnerability discovery in compiled binaries.
> Identifying bugs in code is a core part of the secure software development lifecycle, and unblocking this allows for software engineers and coding hobbyists alike to produce more secure code, reducing new vulnerabilities put out into the world.
Not happy with these annoying "safeguards" but at least it's a step in the right direction. Looks like Opus 5 has the same vulnerability detection performance as Fable 5 and that makes it worth it for code review.
This is excellent model. I was working on some Linux kernel code, and Sonnet 5, Opus 4.8 had given up on the problem i was trying to fix (after several hours). Opus 5 was able to triage and fix the issue in under 30 minutes.
So is it better or worse than Fable 5 on large-scale complexed software engineering tasks (distributed systems, operating systems, high performance computing, compilers, etc. - not frontend stuff)?
This is actually pretty cool. They took all the nice parts of Fable and trimmed the dangerous parts out. I guess distils do work (in some cases). It's very clear they did the same with Sonnet 5, due to the way it orchestrated things, but I think Sonnet was the wrong model to distil Fable from.
It’s funny to share benchmarks showing Opus 5 scoring better than Fable 5 across the board and then saying “but it isn’t actually better than Fable 5”. So then what’s the real definition of better? And why post all these numbers if even you don’t trust them?
It looks like claude-opus-5, likewise most Anthropic models run in Claude Code, sometimes fails to create a TODO list before jumping in to fix a small bug https://github.com/marcindulak/claude-fails-to-follow-claude....
The desire of the models to act at the cost of ignoring user instructions is still noticeable.
It's pretty wild how we are seeing the conversation change every 1-2 weeks. I wonder how long this cycle of progress and innovation among competitors can keep up.
the writing style is so bad. cryptic sorta smart sentences that beat around the bush. say what you are trying to say!
Anyone else feel like Claude Code has gotten worse lately? It keeps going off on tangents I never asked about, and it won't stick to a simple rule I've given it repeatedly: stay concise, only expand when I ask. It just doesn't follow that.
Worse, about two weeks ago it recommended a command and assured me it was safe. I pushed back and asked it to double-check, and it confirmed again that it was safe. I trusted that and ran it — and it wiped out weeks of my data.
I don't have hard proof, but I can't shake the feeling that Anthropic is doing what Apple does: rolling out a new model while letting the old one degrade, whether on purpose or just as a side effect (like how iOS updates quietly eat more resources and slow down older phones). Lately I feel like I'm constantly fighting with Opus (not Fable, because my Fable quota burns through way too fast).
For anyone wanting a faster overview, I used NotebookLM to create a brief video summary after going through the system card and announcement blog using a cinematic video overview. Link: https://www.youtube.com/watch?v=SUFBhvQ2tY4. And a podcast companion: https://www.youtube.com/watch?v=nYZTW2snXow
Opus 5 Pelican SVG: https://pelocan.ai/drawings/ese0s599
Interesting, they finally support `system` messages anywhere in a chat conversation:
> Mid-conversation system messages are available on the Claude API, Claude in Amazon Bedrock, and Google Cloud. > > This feature is available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5. No beta header is required. This feature is not available on Claude Sonnet 5; use the top-level system field instead.
For nearly all models EXCEPT Sonnet 5? That is weird. How old is Sonnet 5 really?
After Opus 4.8 intelligence really started to matter less and less for the programming tasks I have. If I have to handheld anyway, why would I wait more or pay more?
Looks like the API price in tokens is same as previous Opus or Sol, double the price of Terra.
Maybe there’s a better comparison than cost per token, but it will be application-specific.
Better than Fable 5 on all but 3 evals.
Has Anthropic ever mentioned how do Opus and Fable differ? It used to be Haiku < Sonnet < Opus in terms of params. Where does Fable fit in this?
Is Fable 5 just Opus 5 with some additional long-context management modifications for extended self-directed work? Or are they actually truly different models?
Why can you actually still use Fable 5 when Opus 5 is half the cost and as good as Fable?
I have started distilling
Using `/model claude-opus-5[1m]` you can also use the model with older versions of Claude Code
"although Opus 5 shows improvements in its ability to identify software vulnerabilities, it is substantially behind Mythos 5 in its ability to exploit them."
"Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels".
This is probably great news, but then again, where does this leave Fable as a choice?
my early and non scientific feeling:
- it has this annoying Opus response style(since Opus 4.7) with bunch of very hard to interpret word salad
- on >xhigh it eats tokens like there is no tomorrow
I don't like it. Since Fable is unaffordable for anything meaningful, I'll stick with Sol for now. I was on Max 5x, saying hi to Fable costs %5 weekly.
the first model that can one shot a proper qix game - i am impressed
https://claude.ai/public/artifacts/3ea4da3e-76b8-4b9e-acd9-3...
As a coder, I’ve had no desire to use Fable. In fact I switched from Opus models to sonnet 5 and haven’t noticed any drop in quality on large repos. It seems the gap at the top is very small and not hugely noticeable for backed/frontend. Has anyone else had this experience?
Soo most of the benchmarks are better than fable... Is this naming scheme just to avoid getting banned again?
I recently had Sol spending 40 USD circling around a simple task, recently in real world coding Anthropic seems to be ahead a bit.
FYI: `/model claude-opus-5` works to use it even through `/model` still tries to serve 4.8
I found opus 4.8 too agreeable and too wordy(as opposed to codex) and too agreeable. If you are reading documents generating by it was too much. TBH. Fable did a bit better on this. Anyone seen a marked difference with opus 5 on this?
Given that their chart cost axes are almost always log-scale, I’ve noticed starting with Fable that the Low and Medium effort settings might actually be worth setting as your default.
Nerds should stop giving away their work by open sourcing it, or letting AI see it. All you're doing is helping them create your replacement. Stupid nerds.
Anyone has an insight into how much money labs are putting into benchmarks?
Just Arg-AGI-3 is quoted above 20K USD and footnote says average of 5 runs (!!). Likely just a drop in the bucket to the training budget but still..
Where's the reset...
The benchmark appears to have a mistake, as Opus 5 and Fable 5 score 53.4% and 53.5%, respectively, for the Agentic Coding row (FrontierCode v1.1). But Opus 5 is the highlight.
What really impress me is opus 5 is better in alignment than fable 5!
Very interesting to see such a focus on cost for performance here
Very impressive headline benchmark numbers. I expected a step change, but not past Fable. That said - it all depends on whether the classifiers make the model unusable...
Same cost as 4.8 but better that 4.8. Happy to get more efficient model. But is there any reason all companies are releasing models back to back after GLM 5.2.
It starts at page 148.
The internally-reported benchmarks (Frontier-Bench, AutomationBench) and the customer quotes (Cursor, Devin, Lovable) all have a commercial stake in the outcome
worth waiting for independent evals before drawing conclusions.