logoalt Hacker News

shwajyesterday at 5:19 PM13 repliesview on HN

I like how they highlighted Opus 5 as the best for “Agentic Coding” even though the number is slightly lower than Fable. Close enough for marketing, I guess!


Replies

dd8601fnyesterday at 5:22 PM

At half the price and less likely to auto-downgrade, it sounds like a reasonable claim.

show 2 replies
ceejayozyesterday at 5:21 PM

Best can describe multiple things.

Almost as good for half the cost is something I'm very comfortable describing that way.

show 2 replies
tshaddoxyesterday at 5:28 PM

The blog posts figure cites Frontier-Bench for its agentic coding score, and shows Opus 5 beating Fable 5 43.3% to 33.7%.

ActivePatternyesterday at 5:29 PM

I think you're being overly cynical here. First, I don't see any claim that is the world's best model for agentic coding. Second, it is absolutely the best model in terms of coding performance vs. dollar, and it's raw performance seems very close to the frontier.

show 2 replies
manojldsyesterday at 5:32 PM

Which numbers are you seeing? It does show that it's better than Fable 5 in most things related to coding?

Aurornisyesterday at 5:26 PM

Using the most expensive model for all of your agentic coding work hasn’t been good practice for a long time. Not unless you have infinite money to spend.

Fable is typically used for key planning, architecting, and review tasks.

I think this is a case where you don’t understand the use case, not that the marketing department is making mistakes.

show 2 replies
jsLavaGoatyesterday at 5:36 PM

In my opinion, the frontier is passed what is really needed for coding. Fable is good as a supervisor.

throw03172019today at 3:19 AM

And no data retention for 30 days.

toephu2yesterday at 6:25 PM

Also it scored worse on DeepSWE than chatgpt 5.6 sol

dbbkyesterday at 6:21 PM

Yeah I spotted this immediately too. I'm sorry. You're supposed to be a multi billion dollar company and you can't even highlight your chart honestly?

unclebucknastyyesterday at 6:22 PM

Recent releases have said something to the effect (paraphrasing here):

"Use <less expensive or older model> for everyday tasks and <other non-critical stuff>. Use <more expensive or recent model> for complex coding tasks, refactoring large code bases, etc.".

Then, the next model/release emerges and the previous "best for complex" gets demoted to "everyday".

Obviously, it's all relative. But, it does beg the question: was the previous model really good for complex coding tasks or no? I mean, how is it now suddenly only good for the "easy" stuff?

show 1 reply
entropicdrifteryesterday at 5:22 PM

I mean that certainly makes it best-in-class