logoalt Hacker News

wkcheng • today at 6:15 PM • 10 replies • view on HN

The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?

Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.


Replies

usaar333 • today at 6:36 PM

Per the charts, there is largely no point to using Sonnet 5.5 at high+ as opus low generally will give similar performance at similar or lower cost.

But Sonnet 5.5 at medium and below gives you a cheaper option at a performance worse than the lowest thinking Opus (low), which may be viable for "low intelligence" use cases.

➕ show 1 reply
ricardobeat • today at 6:30 PM

At low and medium effort it is 1/3 cheaper, at high it’s a step above Opus/low. It only looks worse at xhigh.

➕ show 1 reply
delillos • today at 6:31 PM

There's a sort of magical thinking needed to answer a question like that. You might say it comes down to "feel" of the model; i.e., the indefinable differences in the way that they speak to the user and approach problem solving. Perhaps Opus is suited for tasks that tackle new ground, while Sonnet might be better at tasks that are more grounded in the code.

Ultimately it's slightly ridiculous to define model capability on a single axis. It's like a standardized test. Sure, you can line people up by their ACT score, but that doesn't mean a doctor and a brilliant artist who both do well on the ACT have an identical intelligence or approach to life. It just can't be captured.

benjiro29 • today at 8:31 PM

The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?

This screams to be that Sol vs Terra model problem that OpenAI had. On paper half the price, in actual usage the price gap was so close for less good results, that everybody just spammed Sol.

RussianCow • today at 6:31 PM

It appears, at least from a quick look, to be noticeably faster than Opus. If true, and you don't need xhigh/max reasoning for your use case (like a well-defined set of code changes), Sonnet might get the job done much more quickly.

With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.

Jcampuzano2 • today at 6:41 PM

I'm honestly not sure where they're getting their 30% numbers from at all. In every single chart that they chose to display except for one, it costs similar or more than Sonnet 5, while also being comparable in price to Opus.

Maybe it's buried within their system card but I think that this would be one of the first things they'd want to show in the announcement article and they fail to do so.

I really don't know who does Anthropic's marketing but they always seem to a pretty terrible job in their announcements from my perspective.

dominotw • today at 6:53 PM

just shows you how little control of output these labs actually have. They are training two models that kind of ended being the same so whatever they were doing specifically didnt make much difference.

SubiculumCode • today at 6:26 PM

t/s maybe? IDK, because their token speed comparison was against Sonnet 5.

solenoid0937 • today at 6:31 PM

It literally does not?

quatotor • today at 6:21 PM

[dead]