logoalt Hacker News

Sonnet 5.5

464 points • by D2OQZG8l5BI1S06 • today at 5:58 PM • 314 comments • view on HN

Comments

simonw • today at 6:44 PM

Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.

https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Here's how the thinking effort levels compare:

  low
  27 input, 1,623 output, thinking_tokens: 0
  1.6284
  Duration: 10138ms (10s)
  
  medium
  27 input, 1,796 output, thinking_tokens: 0
  1.7914 cents
  Duration: 11266ms (11s)

  high
  27 input, 2,334 output, thinking_tokens: 745
  2.3394 cents
  Duration: 17376ms (17s)

  xhigh
  27 input, 5,730 output, thinking_tokens: 2535
  5.7354 cents
  Duration: 41882ms (41s)

  max (failed to return response)
  27 input, 128,000 output, thinking_tokens: 128000
  $1.28
  Duration: 940617ms (15m 40s)
Low and medium both used 0 thinking tokens.
➕ show 11 replies
Sol- • today at 6:13 PM

Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.

More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).

So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.

So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.

Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.

➕ show 16 replies
saint-evan • today at 10:03 PM

unwittingly said 'yaaay' when I saw the Sonnet 5.5 entry. lol I love Sonnet so much. Loved it since 3.5 and never really liked Opus even when I tried to use it for technical work. Once we had two back to back anthropic releases neither of which was a Sonnet upgrade from 4.5 (4.6?) and I was getting kinda sad that they're considering discontinuing it. These are weird reactions I'm having to these tools even when I mostly use them for technical work considering I prefer talking to GPT and Gemini is just a blunt, very powerful hammer.

abejora • today at 6:17 PM

Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.

Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.

[1] Section 8.5 of the Sonnet 5.5 System Card

➕ show 8 replies
wongarsu • today at 6:11 PM

"Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5

Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models

➕ show 2 replies
MisterMunchkin • today at 7:16 PM

It costs 20x more than the Chinese models I use. I just don’t need them anymore. Sure I’d use them if forced to for a job, but I don’t pay them outside of that anymore.

And my job won’t even pay for Claude now because it’s so ruinously expensive.

➕ show 3 replies
wkcheng • today at 6:15 PM

The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?

Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.

➕ show 10 replies
Jcampuzano2 • today at 6:25 PM

I don't understand why I would really use this over using just a lower or even similar effort level on Opus, given that in many of the benchmarks it's basically the same cost, if not more, at any effort higher than medium.

Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.

Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.

ghoshbishakh • today at 6:27 PM

So sonnet is better than Fable now? That Fable which was too dangerous to release? I am so confused now.

➕ show 3 replies
johnmlussier • today at 6:08 PM

Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.

This is bollocks. Their safeguards are shit.

➕ show 16 replies
robertclaus • today at 9:19 PM

These benchmark results keep getting more questionable without error bars.

gregwebs • today at 6:51 PM

This is better priced than Opus for tasks that are token heavy but not complicated. But a quick look shows that at least on some benchmarks DeepSeek performs as well and of course the cost is an order of magnitude less.

From looking at their Terminal-Bench graph, anything you would use level "high" or above for Sonnet it seems like you should consider using Opus instead.

OpenAI Luna is a lot cheaper. But DeepSeek seems smarter and the cost seems similar.

guilhermeasper • today at 8:41 PM

AI companies these days releasing models every week like Netflix episodes.

➕ show 1 reply
heyjstn • today at 6:49 PM

Have anyone tried a workflow that:

- Fable 5.1 for planning/adversarial reviewer

- Opus 5.5 for well-scoped tasks break down

- Sonnet 5.5 for these well-scoped tasks implementation

I think the blocker might be how efficient the context is compacted and sending around between these agents

➕ show 3 replies
nanook • today at 7:38 PM

Sonnet is 1/5th the price and seemingly more powerful than fable (the model that was too powerful to release). I can't make sense of this. Why would anyone use fable now? Or are the benchmarks completely pointless and one has to just try em to get a feel for what they can and can't do?

➕ show 2 replies
avree • today at 6:14 PM

Crazy bad front-end design. Site hijacks my gestures so I can't swipe back anymore, starts with a full page autoplaying video...

➕ show 1 reply
yapfrog • today at 6:29 PM

From the graph it looks like I'd rather use Opus 5.5 High than Sonnet 5.5 at all

alansaber • today at 6:31 PM

Always key to include the one bench where the smaller model inexplicably outperforms the larger model

mchusma • today at 8:55 PM

I feel like sonnet is priced too close to opus right now. If Sonnet 5.5 were half its current price it would make sense to use. At its current prices, I won't use it in applications (I would use cheaper models) and I won't use it in my subscriptions ( just use Opus instead). At least that is my initial reaction.

dom96 • today at 7:05 PM

I built an adversarial esoteric programming language to benchmark LLM models and just ran it on Sonnet 5.5 It does worse than Sonnet 5. Mainly because it is more reluctant to keep going to get an answer, instead it returns to ask the user questions whether to keep going.

https://bench.killswitch-lang.org/

    Claude Sonnet 5    17.8%
    Claude Sonnet 5.5  7.4%
onlyrealcuzzo • today at 6:14 PM

> In our testing, it costs up to 30% less per task than its predecessor.

> Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.

This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.

They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally a year behind.

Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.

I've almost exclusively been using Anthropic for design and review, as it almost never makes sense to use any of their models for implementation (90%+ token usage) - except in the rare cases it's something too complex for a number of 10-100x cheaper models (and more importantly for me 5-10x faster, too).

For me, it's less about cost. I'm not doing anything that can't be done with a $200 subscription and minimal intelligence on what models to use. It's primarily about speed. I don't have an entire work day to give Opus / Sonnet a task that Flash can get done 95% as good in 30m.

This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.

Hopefully they release a Haiku that actually has a reason for existing.

➕ show 3 replies
sergdigon • today at 9:38 PM

Am I getting out of touch or is it becoming kind of confusing what model should be used when? Sure you have tons of benchmarks pareto cost/perf curves etc but at the end of the day when I have a task to give to a model it is not so clear which model and which effort I should choose ... Also benchmark numbers are often reported with max effort but by default effort is medium and based on the pareto curve on this page, Sonnet 5.5 seems more cost efficient than opus only if effort is low or medium!

AM1010101 • today at 6:50 PM

For me I would like to pair this with Opus 5.5 as orchestrater and use Sonnet as a sub agent. Therefore I want it to be fast when on low or medium and not break the bank.

On low and medium it seems competitive, maybe slightly cheaper than opus, in terms of intelligence per task.

If the time per task is lower (Artificial Analysis don’t have the date up at time of posting) then I have a clear use case for this model all other things being equal.

trvz • today at 8:44 PM

Maybe they could put Fable onto creating a website that doesn't use 80% of the GPU on an Apple M2.

solenoid0937 • today at 6:30 PM

Amazing release. This thread is already full of cynicism and angry hot takes. The Opus 5.5 thread was like this as well despite it being a hit with everyone.

At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.

➕ show 3 replies
tombert • today at 6:39 PM

I like that "alignment on safety" appears to mean, at least for anything I've been doing, that they won't violate Microsoft's terms of service. I even had it pushing back on me activating an LTSC key on Windows because LTSC keys are "often purchased on a gray market and violate Microsoft's TOS".

➕ show 1 reply
sajithdilshan • today at 7:26 PM

I use Claude Code everyday for work and the main model I use is Opus (For planning, breaking down tasks, writing tickets, implementation, etc.) and Haiku for running tests. Honestly have no idea what is the use case for Sonnet

➕ show 1 reply
bastawhiz • today at 8:18 PM

I'm confused by the charts comparing it to Opus 5.5. It looks like slightly lower accuracy for the same cost along most comparisons. Am I reading that right?

Is it just the benchmarks? Because otherwise it suggests it's twice as chatty as Opus for a comparable output... Which kind of defeats the purpose

alasano • today at 6:42 PM

I wonder if Fable 5.5 is coming this week to drown out the OpenAI dev day announcements

ChickeNES • today at 6:20 PM

Weirdly, the web ui has Sonnet 5.5 as "Most efficient" for "simpler tasks" and 5.0 still labeled the same for "everyday tasks", with Opus 5.5 as "For complex work and everyday tasks".

s3p • today at 6:10 PM

I'm loving the tit for tat cost charts these guys are doing. Just a few days ago it looked like OpenAI ruled the cost pareto frontier. Not even a week later and Anthropic is taking the charts again. See you guys same time next week?

➕ show 1 reply
low_tech_punk • today at 8:18 PM

The documentation mentions error code "frontier_llm": The request could assist the development of competing AI models.

I'm very curious how do they know what requests could assist competing AI models.

croemer • today at 6:40 PM

Playing around with it for a few minutes, Sonnet 5.5 feels very fast, much quicker than Opus 5.5. Can't tell yet if it's a lot worse but the speed is definitely welcome.

swingboy • today at 7:07 PM

Is Opus still 2x usage of Sonnet after this? My Claude Code isn't showing that warning anymore when I look at /model.

ghoshbishakh • today at 6:33 PM

So Sonnet 5.5 on max effort is as expensive as Fable 5.1? Because it uses a ton of tokens for a task.

In xhigh effort it is a lot cheaper and possibly lot less impressive?

bayesianbot • today at 6:20 PM

Cache reads priced the same as Opus 5.5? So there won't be that much price difference in agentic coding. Or is that a mistake in the table, that seems quite weird

pookieinc • today at 6:08 PM

It's interesting that in all their benchmarks, they omit Fable numbers and only focus on Opus, Sonnet, and OpenAI models. Maybe Fable is out the door?

➕ show 5 replies
itishappy • today at 8:05 PM

Wow, I've never seen a site break chrome this badly. I get a black screen then it stops rendering the entire window, even when opened in the background.

s314 • today at 6:36 PM

In the Artificial Analysis Intelligence Index, Claude Sonnet 5.5 is the second best model behind Opus 5.5. This however is with max effort which costs even more than Opus 5.5 max. But Sonnet 5.5 xhigh is cheaper than Opus 5.5 xigh and matches GPT 6 Astra xhigh in the benchmark.

➕ show 1 reply
taurath • today at 6:12 PM

After 5.0 I feel the need to give a long eval period before deploying it with enthusiasm as I did with 4.6 which felt like a big leap. Codebases all through my company which is very seem to have taken a dive in quality, with nonsensical and unreadable multi-line comments wherever devs are letting the models run free.

➕ show 1 reply
a13o • today at 7:47 PM

This doesn’t have an interesting footprint on the intelligence/cost Pareto line compared to existing Opus 5.5 and GPT-6 models.

__jl__ • today at 7:51 PM

artificialanalysis.ai benchmarks are [here](https://artificialanalysis.ai/articles/claude-sonnet-5-5). Anthropic is back at spot 1, 2, 3 and 5. Impressive even if these benchmarks are problematic in many ways.

➕ show 1 reply
mroche • today at 7:22 PM

Is there ever any focus on producing new Haiku models? There are a lot of use cases for quick to return models when you're limited to a single provider.

takerofnaps • today at 6:11 PM

Sonnet 5 seemed somewhat benchmaxxed to me. So was Opus 5. I wonder if this will be as big of an improvement as opus 5 -> opus 5.5. Maybe I will switch back from GLM 5.3 flash for some tasks.

pavitheran • today at 6:24 PM

Big jump on Agentic coding from 10.3% -> 70.6% from Sonnet 5 -> 5.5 which even surpasses Opus 5.5. Opus 5.5 is really strong so this is impressive especially for the cost.

➕ show 1 reply
square_usual • today at 6:25 PM

Once again, once you hit the high/xhigh level you're better off using Opus low/medium to get better results for around the same price. So I suppose the main point of this release is that you have a lower end than Opus low, which I suppose some people will like?

➕ show 2 replies
rtuin • today at 6:14 PM

Any benchmarks other than computer use/agentic coding published yet? Curious to compare more broadly with other models

ramish94 • today at 5:59 PM

In terms of benchmarks for agentic coding, it basically stacks up nearly 1:1 with Opus 5.5.

Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)

FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)

CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)

Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family

➕ show 3 replies
_fw • today at 6:16 PM

I still can’t find a place for Sonnet models, I never have.

I bounce between ”fuck you, give me an AGI-approximate robot god” or ”how dare you charge me more than $0.04/million tokens”.

Give me the frontier, or give me the cheapest form of good enough.

➕ show 2 replies
limsungkee • today at 6:33 PM

Yesterday, I realized that Opus 5.5 is cheaper than Sonnet 5. Now I know the reason.

🔗 View 15 more comments