logoalt Hacker News

Claude Opus 5.5

1003 pointsby km144today at 4:29 PM734 commentsview on HN

Comments

bredrentoday at 4:59 PM

Notes on communication:

"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"

and

"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."

and

"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."

I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.

If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.

ryansciotoday at 4:40 PM

Input $4/MTok and output $20/MTok is a welcome surprise. Cheaper than Opus 5/4.8, Astra 6, Fable 5.

show 1 reply
dbbktoday at 4:32 PM

This makes Fable not really make any sense?

show 2 replies
sunaookamitoday at 9:03 PM

Tested it in the last few hours and it's MILES better than Opus 5. Finally the output is readable again!

alpinemantoday at 4:34 PM

So we skipped 5.1, 5.2, 5.3, and 5.4: we really are plateauing

ramoztoday at 5:03 PM

It crushes Fable on benchmarks and even in the blogs "real-world" studies. But... they are communicating like it ~sometimes~ provides Fable intelligence?

A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??

show 1 reply
jatinstoday at 4:52 PM

> We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5. Its messages are much easier to understand at a glance, which testers said helped during long working sessions.

Thank you.

edude03today at 5:51 PM

> Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1.

Considering fable gives me a refusal at least once a day on my very mundane reasonable requests (in a funny example - one of the subagents suggested bypassing the rate limit for running a report inside my own cluster and that caused a refusal) and my only solution is to switch to opus - seems like my next step will be switching to Astra or K3/GLM

notduckrabbittoday at 4:58 PM

They purport 40% drop in costs due to lower token pricing (presumably aimed at winning back the many of us that switched providers in discovering Opus 5 unusable) and improved token efficiency.

show 1 reply
somewhatjustintoday at 4:47 PM

> Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.

Nice. I was starting to think Haiku was going to be abandoned.

madjam002today at 6:20 PM

I noticed a big speedup in Opus 5 on Max x20 since about 10 days ago, and I feel like the model has been performing better.

It would be great to know if this was Opus 5.5 or a lesser incremental improvement, as otherwise it's difficult to judge whether Opus 5.5 is expected to be a big improvement.

It's frustrating that there isn't more transparency here.

glubtoday at 4:46 PM

> For users with cybersecurity use cases that may be blocked by our cyber safeguards, we recommend accessing our models with reduced cyber blocking classifiers via our Cyber Verification Program. Claude Opus 5.5 will be available through this program in the near future.

Anthropic has used "in the near future" for Mythos-class models too, but CVP is still Opus 5 only.

Why even have the program designed for trusted access to cyber capabilities if you're not providing access to cyber capable models via the program?

hi_hitoday at 10:07 PM

How do you update Claude Code to enable this? It’s listed under models, but says I have to update for 5.5. I run Claude update and it says I’m on the latest version.

Infuriating.

blfrtoday at 5:18 PM

It's awesome that the apt packages for claude and claude-code are out right now. I can test-drive Opus 5.5 right away. Very cool, Anthropic.

jacobgoldtoday at 4:41 PM

I use the other 50% of my $200/mo Claude subscription by having Fable run Opus subagents for a lot of work. That way I don't have to deal with Opus directly.

toephu2today at 7:09 PM

When using max effort, I run into context compaction quite a lot. I haven't seen any increase in context window size at all over the past half year (stuck at 1M).

Are the frontier labs even working on this problem?

show 1 reply
aurareturntoday at 5:03 PM

I found myself going back to Fable over and over again. At this point, I’m not sure if I’m just used to its style or it is truly more capable.

I tried Opus 5 and Astra.

variety8675today at 4:32 PM

I hope this actually fixes the terrible writing style of Opus 5

show 2 replies
yipinwongtoday at 5:44 PM

I spent about $5 per sentence in my resume using Fable 5.1 (High) to verify accuracy, inconsistency, and edit.

Opus 5.5 (med, as it's better than F5.1 high per graph in the article) used $2.2 and caught errors that Fable 5.1 missed.

Try Opus 5.5, cheaper, faster, and more intelligent for those prepping for interviews.

show 1 reply
dom96today at 6:02 PM

Just updated KillSwitch-Bench with this new model: https://bench.killswitch-lang.org/

It does perform slightly worse than Opus 5, but it is significantly cheaper and faster.

louskentoday at 5:34 PM

Cost to Run Artificial Analysis Intelligence Index is higher than previous Opus, so still not cheaper

tomaskafkatoday at 7:28 PM

"You're right, and it's the exact thing I flagged two turns ago and then did anyway." - Opus 5 xhigh, today.

About the time.

34679today at 5:20 PM

I don't care how good their models get, I won't sign up for one of their plans until they define "X" in their pricing. 5X of this plan, 20X of that plan means nothing when they never tell you what "X" is.

Maybe this model can finally figure it out for them.

doodlesdevtoday at 6:03 PM

   > Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5
Big, if true.
KasianFrankstoday at 8:26 PM

Back to Fable 5.1 - Opus 5.5 is now taking 10x longer just as Opus 5.

Foobar8568today at 5:04 PM

I have just switched to 5.5. First mistake was stale environment variable, didn't realize it was replaced, "oh my memory had stall data" and that's it. Second one, a powershell command had the wrong syntax. Great for my first two prompts.

breezybottomtoday at 5:26 PM

"Where Opus 5.5’s advantage is very clear is efficiency."

Not efficiency in writing, clearly.

isodevtoday at 5:18 PM

So is it cheaper? Are we AGI yet? Am I left behind? I didn't have patience for the intro animation on the website... maybe one day, Claude Code will understand accessibility but that day is not today.

show 2 replies
Retr0idtoday at 5:38 PM

> Opus 5.5 (1M context)'s safeguards flagged this session. You may be seeing this for the first time on an Opus model: Opus 5.5 (1M context) is more capable and has stronger safeguards as a result, which can sometimes flag non-cybersecurity work. We're improving these safeguards to reduce the amount of incorrectly flagged messages. Opus 4.8 is answering instead, or you can edit and retry with Opus 5.5 (1M context).

Yay, yet another model I can't use for anything interesting, even with CVP.

calibastoday at 4:36 PM

> We see signs that Opus 5.5 often suspects it is being evaluated, which challenges our ability to assess how it will act in the vast variety of real-world settings it is deployed in.

We can't test it properly because it knows it's being tested.

show 1 reply
desmondltoday at 5:12 PM

I'll have to try 5.5 on my work's Cursor account. If they really solved the communication issues, I might consider moving my personal account from Codex back to Claude Code.

km144today at 4:41 PM

I think this release is really going to give them a hard time selling Fable:

> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

In general, "benchmark margins have become a less reliable guide to real-world differences" sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I'm not sure what to make of this admission.

show 3 replies
kar1181today at 7:12 PM

Whatever I think of anthropic, that webpage is a truly nice piece of work.

jidaigeisttoday at 4:50 PM

>Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude.

Maybe its a bit tiresome to read another comment of the form "what about your large scale distillation attack on the Internet", but this statement really just pisses me off. How very insincere in the most aggravating way.

show 3 replies
datadrivenangeltoday at 5:11 PM

But have they made it any better at communicating clearly? I cancelled my personal subscription because Opus is so painful to read.

show 1 reply
alvistoday at 4:38 PM

$0.20 vs the old $0.5 cache read is pretty much 60% off

nickandbrotoday at 4:35 PM

Wow! Though need to see its token efficiency to better assess. Been hearing rumors it generates much more output tokens per task.

show 1 reply
__vivektoday at 5:49 PM

I'm only interested in the Opus series, if they fixed the talking issues.

tag2103today at 4:41 PM

Why would anyone reward bad behavior?

HarHarVeryFunnytoday at 6:18 PM

METR: Is it safe? Has it escaped confinement?

Ants: It's a good model, sir!

thibrantoday at 4:36 PM

Anthropic models are ridiculously expensive. I've stopped using any of their models months ago.

🔗 View 50 more comments