logoalt Hacker News

throwaw12today at 7:23 AM10 repliesview on HN

Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)


Replies

OtherShrezzingtoday at 8:11 AM

It could be that the set of your day-to-day workload which could feasibly be accelerated by AI just happens to be saturated around Opus4.5, but you can still see lots of “reasoning” which makes you think the model is more performant in the first days of use. That’d mean you couldn’t perceive any meaningful difference in more powerful models’ results, even though you can see a difference in the raw output due to the length of reasoning traces leading up to the result.

So for example, if your workload was literally just addition of sets of numbers, you’d never have noticed progress in the result beyond GPT3.x level models. But you would perceive a difference in the now-Tolstoyan length reasoning text accompanying the result.

SubiculumCodetoday at 9:02 AM

5 seems incredibly smart to me in my conversations today about some pretty niche ideas in.longitudinal modeling. It.felt.like a big step up.from 4.8, to me

show 1 reply
staredtoday at 7:43 AM

It's called frog boiling.

We get used to the new level of intelligence so fast, any deviation feels like going back to the stone age.

If you don't believe me, create something complex with Opus 5 and then with Opus 4.5, and notice the difference.

show 1 reply
yorwbatoday at 7:40 AM

Well, what kinds of things do you see Opus 4.5 completely fail at? Maybe those are not the ones that newer models have improved on.

submetatoday at 2:12 PM

Some 20 years ago, the telecommunications sector in Germany was liberalized. Many telephone card providers entered what had previously been a barely competitive market. They advertised their products with aggressive claims like: “Buy our €10 top-up card and get 660 minutes to destination X.”

For the first few weeks, they would actually provide those 660 minutes to establish trust in their cards. But after a while, they would quietly start reducing the number of minutes on subsequent top-ups—say, from 660 minutes down to only 300. They wouldn’t do this for every card, so it was difficult to prove. Instead, they relied on averages across their customer base to make the economics work.

Lately, I’ve found myself wondering whether something similar may be happening with frontier AI models. Companies launch with an exceptionally strong model and generous compute limits to build adoption. Once the model is established as a market leader, the incentives change, and users may start perceiving the service as becoming more constrained or less capable over time.

I don’t have evidence that this is what’s happening with Anthropic—or with any other AI company. It’s simply a pattern that the current situation reminds me of.

tudelotoday at 8:09 AM

I honestly just use GPT models nowadays, Claude models are too restrictive and more of a quitter and fable/whatever is just too expensive to be worth it.

slopinthebagtoday at 3:30 PM

Same, like I prefer 5.3 codex over the “stronger” models.

lwansbroughtoday at 7:29 AM

Going to call it user error if you find Opus 4.5 better than 5, sorry.

rf15today at 7:57 AM

I've worked with these systems for four years now and they have not meaningfully improved in that time frame.

We still have:

- statistical correlation between two things will always cause one thing to lead to the other, no matter how much you prompt it to not have that connection (to be expected with a stochastic system)

- Math completely fails in longer contexts

- "thinking" token generation being on the correct track just to 'no, wait' on an already correct conclusion

- smearing of properties between logically distinct objects (a red ball and a green cube can quickly become a red cube and a green ball)

show 5 replies
sscaryterrytoday at 7:41 AM

Enshittification.