logoalt Hacker News

onomojotoday at 7:30 PM9 repliesview on HN

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.


Replies

cromkatoday at 7:42 PM

Agreed, it's extremely frustrating. It's the only model that actually makes me curse when talking to it, even knowing how counterproductive it is.

show 2 replies
Fordectoday at 8:35 PM

Yeah, I've dropped back to 4.8 entirely for the remainder of this billing cycle. I'm going to be seriously looking into Qwen adoption and harness migration options over the course of August.

copperxtoday at 7:35 PM

I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.

show 2 replies
visargatoday at 7:38 PM

Sent to solve one task, came back with half of it solved and 2 more problems.

show 1 reply
enraged_cameltoday at 8:08 PM

It's my daily driver. I like it and find it noticeably better than Opus 4.8.

After I started reading complaints about Opus 5, I gave Fable the task of evaluating a bunch of code Opus 4.8 had written and compare it to Opus 5's code. Fable ran a dynamic workflow and the scores came back 15-20% higher for Opus 5's code in terms of quality, correctness and readability/conciseness. I did not tell Fable which Opus wrote which code, and I turned off memory as well to ensure there was no pollution from that angle.

My only complaint is that Opus 5's prose is annoying as hell. I wrote a custom skill for it for concise debriefs and it has been working pretty well for me.

nomeltoday at 8:07 PM

What's the clear best, that you see?

show 3 replies
fellowniusmonktoday at 8:24 PM

I have some internal tests I use for areas where one particular solution/paradigm is dominant but worse.

Opus 4.6 is the last model that's actually useful and can "adjust" its perspective to use the newer & better solution.

Where Opus 4.8-5 has over fit training on worse/older but "dominant" solutions it refuses to adjust.

Not only does this create an existential threat to adopting progress but it also means that if you have a code base that has rare but real world tradeoff the newest versions of Opus 4.7, 4.8 and 5 are worse than useless and become a major dev timesink.

logicchainstoday at 7:40 PM

"As you requested, I've finished task X. Honestly, task X turned out to require task Y, which I haven't actually done. Task Y is the next step if you'd like to continue along this route."

show 1 reply
bontaqtoday at 7:48 PM

It's an infuriating model