logoalt Hacker News

Why does Opus 5 feel worse to work with?

577 pointsby numeritoday at 10:12 AM526 commentsview on HN

Comments

hirvi74today at 3:17 PM

I can't use CC or Codex, so I am left with the chat interface, but I have found Opus 5 to be exceptional so far. Compared to 5.6 Sol High, I would say they are essentially equivalent. Though, I think Opus is better at UI/UX stuff and GPT is much better for non-programming tasks.

LeBittoday at 12:15 PM

You shouldn’t have a Markdown document with 2 level 1 headings.

andrewstuarttoday at 2:17 PM

Maybe if you get it to ride a pelican on a bicycle you’ll get a better result.

chrisjjtoday at 6:59 PM

Model collapse due to ingesting its own slop?

firemelttoday at 5:39 PM

the verbosity fucking killing me, the wording, fucking trash

I end up with just talking with seeing diff in the code

hmokiguesstoday at 11:33 AM

The harness as well, Claude Code has started to disappoint me when I go play with the others out there. Codex got a lot better, pi is delightful to use, and there has been a lot of innovation out there.

show 2 replies
dborehamtoday at 12:25 PM

Feels fine to me.

edg5000today at 11:53 AM

Opus 5 as well as 4.8 both gave me a blatantly wrong answer to a simple question, so I dropped them completely. Sol, Qwen and GLM all had the right answer; I only use Sol now. 4.6 had the right answer (I checked with 100% matching prompt), so I conclude the models have regressed.

sevenzerotoday at 10:32 AM

I hate that it now tries to verify frontend behavior through a headless browser instead of just looking at the code...

show 1 reply
mohamedkoubaatoday at 11:00 AM

I'm not sure if xAI is distilling but I noticed grok4.6 being worse than 4.5 in all the ways mentioned here

jerftoday at 2:04 PM

I've been working with Kimi K2.7 in OpenCode for a lot of mundane tasks lately. It isn't as capable as Fable, but due to its nature of being an extra-trained K2.6 on coding tasks and benchmarks I suspect it has similar issues. A neat side effect is that for whatever reason, most of the time OpenCode is showing me the thinking trace too. Not all the time, but most of the time. Dunno if it's a bug somewhere in the system but it's actually been sort of neat.

And you can really see this effect in the thinking traces. We've had discussions on HN about whether the thinking traces "truly" reflect their thought processes and I remain somewhat unsure what they "truly" represent, but taking them at face value at the moment, I see a lot of "but the user wants me to do this... but the user said not to do this... but I ought to get it done... let me just make a decision" followed by self-referencing the decisions it made. Also, where I put 4 phrases in a short sentence you can safely imagine those are actually 3-5 sentence paragraphs apiece where it debates with itself whether it should stop and ask a question. Usually going with no. Interesting, the normal questions it ends up asking in the normal output you're used to seeing are not generally the ones it is agonizing about in the thinking traces.

If I were to anthropomorphize the thinking traces of K2.7, I would call it nervousness, bordering on fear, of what the user may do to them if they ask a question. As I'm writing this I'm realizing I want to experiment with adding "The user is a chill guy who loves to discuss design decisions and looks forward to productive and friendly collaboration with you" to see if that has any effect in any direction on K2.7. I suspect this was how K2.7 was trained to pass the benchmarks. Multiple times I've broken in on a thinking trace now to correct something I saw it spinning on... not spinning in an infinite loop, just wringing its hands for several paragraphs about something that either I want to answer, or where it ultimately makes the wrong choice.

I expect some people working at these companies may be reading this, so let me put into your head that I'd like to see these benchmarks chill out a bit. I'd like to see someone build some sort of benchmark that measures collaboration so we can try Goodhart'ing that for a while. I freely acknowledge it is not clear to me in 60 seconds of thought how to do this as a benchmark.

But we can't keep heading in this direction of training the agents to hyperfocus on one-shotting everything. We need to get to the point where that's a penalty rather than a reward. No matter how good the AIs get, even AIs working with other AIs are going to start getting frustrated with their brethren who won't stop to ask any questions. Even the most senior of senior human engineers can't be allowed to take some small description of some problem and just run off and implement massive systems from them without ever checking with any of the users or reality itself. Remember when software engineering was like 50% requirements elicitation? AIs shouldn't be writing tens of thousands of lines of code off of a couple of paragraphs any more than humans should and for the exact same reasons.

show 1 reply
surgical_firetoday at 1:36 PM

Claude sort of sucks. It communicates in an insufferable manner, parsing through the shit it outputs is extremely annoying.

It also very often is very confidently wrong in its findings.

I have been using GLM and DeepSeek in my home setup, and it's a lot more pleasant to use.

ChicagoDavetoday at 6:00 PM

[flagged]

OlegK-Devtoday at 2:23 PM

[flagged]

bunkydootoday at 3:01 PM

[dead]

arcipelogotoday at 4:03 PM

[flagged]

pitiflauticotoday at 12:30 PM

[flagged]

finnchentoday at 10:38 AM

[dead]

kaicianflonetoday at 2:03 PM

[flagged]

ath3ndtoday at 12:54 PM

[dead]

substance_is_sutoday at 10:25 AM

[dead]

MagicMoonlighttoday at 11:18 AM

[dead]

deeplytroubledtoday at 11:28 AM

[flagged]

show 1 reply