logoalt Hacker News

Getting the most out of Opus 5.5 in Claude and Claude Code

114 points • by saikatsg • today at 6:29 PM • 79 comments • view on HN

Comments

rdli • today at 8:49 PM

It’s a really good model. Over the past few days, I give Opus some general directives to basically speed up our CI, and telling it I care both about billing minutes and wall clock time. I told it to create a plan after analyzing everything in our CI, run the plan by a Fable subagent, and then focus on low-risk, high-reward changes.

9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.

➕ show 9 replies
jjcm • today at 10:09 PM

It’s extremely good at frontend, particularly if it has an image reference. I chucked in design reference images into this and told it to focus on the flowing svgs, and it crushed this Star Trek computer-inspired layout: https://html.non.io/lcars-opus-5.5

adastra22 • today at 10:29 PM

Some of this advice is really missing the mark. I will speak to just one I know well. Many of my frequently used prompts have “think through this step by step” because if you don’t, it only considers the task holistically rather than step by step, and different issues emerge in that frame of thinking. I see this. Often when doing planning, for example, it will not notice interdependencies between tasks until you force it to think through doing the whole thing step by step (task by task) then it will notice that step 2 requires a feature introduced by step 14. It wouldn’t notice otherwise.

Yes this has held true on Opus 5.5. I checked. It’s a massively better model, peer to Fable but with different strengths and weaknesses. But it still has this issue. Which to be fair, people do too. Planning is a learned skill.

I think what they’re saying is that the harness no longer uses a text search on “think” to engage reasoning modes. Fair, that’s good to know. That doesn’t mean asking the model to think a certain way doesn’t have the intended effect.

hibikir • today at 9:37 PM

It's much better than 5, but I've had a couple of situations this week where it was too interested in being independent, making calls that went directly against my recommendations. It can also do fun things like convince auto-mode to go way past what I have autorized. For instance, specific permission to run process X in region abz-1 suddenly became running X in 5 other regions, with no warning, and doing modifications that it never mentioned in the summaries. And a few of the times it got the calls very wrong, by assuming it understood systems it didn't. It'd even argue with me when corrected, as it assumed similar names were referring to the same thing, when they weren't.

So asking it to do things on its own for a long time? Given last week, absolutely not.

➕ show 1 reply
jampekka • today at 10:18 PM

What is this spam of generic comments about how Opus 5.5 is so great, with perhaps some anecdote? How is that discussing the submission?

skybrian • today at 10:30 PM

I'm using OpenAI and a different harness, but I'm getting a lot of mileage out of asking "what are the commits?" and "please go ahead, using a Luna subagent for each commit."

I used to have the AI write a planning note with checklists, but this seems good enough nowadays.

satvikpendem • today at 10:28 PM

How are the limits compared to OpenAI models like 6.1 Sol? After the 200 dollar plan rugpull not sure if I should switch, however Anthropic has historically had worse usage limits than OpenAI.

➕ show 1 reply
magicalhippo • today at 9:40 PM

Been very impressed with my most recent project. I wanted to simulate some older electronics circuits. I handed it a folder with scans of old service manuals which contained circuit diagrams. It managed to correctly interpret the circuits, including figuring out some were the same topology despite the diagrams being quite different, or some that had some subtle but very important differences despite looking almost identical at a glance.

In a few cases it asked me to check some subcircuits and some component values because it couldn't read it right. So instead of just making things up it deferred to me.

It also ran tons of small simulation experiments while doing this to verify claims from the service manual, like that the RC filter it had read off the schematics actually had a cutoff frequency that was sensible in relation to some bandwidth number in the manual.

I had uploaded datasheet PDFs for many of the ICs and it used those to cross-reference and validate.

It kept on working for over an hour. When it asked for the manual verification, I described circuit connections in words, like "from pin 3 on IC 2 there's a series resistor of 3k in parallel with a 10 pF capacitor, it then connects to a 18k resistor to ground, a reverse-biased diode to ground, and then finally into pin 6 of IC 4", and it correctly understood the topology in all the cases. Sometimes it asked me to check again because it though something was off, and indeed I had mis-read the schematics.

I also provided reference articles on the underlying theory. Scannded stuff from the 40s and 50s. It correctly read the equations and cross-validated them across papers, and even caught several typos along the way.

I barely had to do anything apart from providing the PDFs and some occasional manual schematic interpretation.

Claude 5.5 on High. Burned through about 50% of my weekly $20 subscription usage, but I didn't try to optimize much.

I did use Sonnet 5.5 Medium on some datasheets and it also did very well on the extraction, but did have to correct itself more often on the conclusions.

➕ show 1 reply
kingcauchy • today at 9:55 PM

I've had troubles with it getting stuck "waiting" for a day on some hook or something in CC that never completed and during a task that was waiting on an orphaned process. That maybe saves token money on checks waiting for long-running processes but makes it hard to trust for long-horizon work.

It's been amazing at making sure OOMs for multiple heavy builds on my machine don't happen, adding queues and locks to make sure performance measurements are isolated and gpu stays clean during experiments.

It's also way more able to execute subagent tasks all at once than GPT 6.1 I tried to give it 10 different subtasks all at once that were overlapping and unrelated issues and it did a good job spinning up isolated worktees, agents and then coordinating the merge back together and then verifying them with agents in batches.

ToJans • today at 9:20 PM

Superb model indeed.

I've given it some big tasks and asked it to parallelize as much as possible etc.

It did burn through my weekly tokens in about a day (20x max), but the output was completely on point. (I knew there was a "reset token usage - opus 5.5" button in my account.)

I've now come to a point where I even delegate my discovery for new features to it.

You still need to give it methodologies though to get the proper output, but the outcome is way beyond what I would be able to realize with a team of 5 in a month.

TomGarden • today at 9:51 PM

Incredible model. I don't see why they can't just include the recommended workstyle as a guided approach into the claude code harness though, and let the people who want to diverge just ignore it

➕ show 1 reply
RGS1811 • today at 10:07 PM

I’m going to be the lone dissenting voice here and say I still find it overeager and irritating to work with.

alansaber • today at 9:15 PM

"Don’t ask it to show its reasoning in the reply" “Explain why you chose this approach in three sentences” says it all really

alwinaugustin • today at 10:01 PM

Can Anthropic introduce a $50 USD Plan ? I dont want to spend 100$ , but $20 plan is not enough for me.

➕ show 4 replies
briga • today at 9:37 PM

Is it still necessary to ask Claude to spin up sub-agents? If Opus 5.5 decides how carefully it needs to think after each question, surely it can also decide whether it needs to spin up sub-agents? GPT 6 series models at least seem to do this agent management automatically.

epistasis • today at 9:39 PM

I really hate long tasks. Claude never gets things right, at least for me, and wastes tons of time when a simple question would have gotten me to the right result rather than several turns of correcting bad decisions in addition to the long amounts of wasted thinking time.

What sort of workloads do well with these long tasks? The big labs are optimizing for long run time on their own, but it seems like a terrible thing to optimize on unless you're trying to do something like prove a hard math theorem, which success is clearly defined and the route doesn't matter a ton.

Plan mode has been made increasingly useless. I need to discuss to iterate to get the desired design, explore options, because Claude never gets it right first try and I don't have enough knowledge of options to specify everything up front.

Ah well, the Chinese models will still work well, I guess.

➕ show 2 replies
danbrooks • today at 9:06 PM

Agreed on Opus 5.5 being a great model. It's the first one that I trust for long running (>1 hour) tasks.

➕ show 2 replies
nozzlegear • today at 10:16 PM

Honey, it's time for your daily Anthropic PR piece!

tebrun • today at 9:19 PM

Opus 5.5 is great and cheaper if you compare to fable with close quality in coding (tested in refactoring java to nodejs), but I do not understand why the week before the release Opus 5 started hallucinating (long loop and waste of token for single tasks)

ares623 • today at 10:29 PM

Is it just me, or is the first example for "define what DONE means" a lot like the "draw the rest of the fucking owl" meme? Except for a small subset of tasks that are already trivial like migration work.

What I mean is, for most tasks I do that aren't trivial, by the time I have defined what DONE means, I would've already did the work and walked the path to get there, which is what I would've hoped to not have to do in the first place.

sergiotapia • today at 9:32 PM

The best model I've ever used easily. It's incredible. I've done so much in the past week. About three months of work I reckon.

➕ show 2 replies
Handy-Man • today at 8:40 PM

Phenomenal model, not sure what they did, but I have been able to do so much with my $20 plan!

➕ show 3 replies
franze • today at 8:40 PM

[flagged]

i_love_retros • today at 8:45 PM

[flagged]