logoalt Hacker News

GodelNumberingyesterday at 6:33 PM14 repliesview on HN

The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M).

This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general.

Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is hard to see ANY improvement:

Terminal-Bench 4.0: Fable 5.1 is +3.5% vs Opus 5.

GDPval-AA v2: +1.5% vs Opus 5.

OSWorld 2.0: +2.5% vs Opus 5.

Humanity's Last Exam (with tools): +1.6%

Keep in mind that this is supposed to be an entirely higher tier of a model than Opus 5. For one tier up and one version up, these are not really improvements. Probably leaves no room to place Opus 5.1 anywhere. Combined with the fact that they are selling 'readability'... Has frontier progress finally stalled?


Replies

johnsmith1840yesterday at 7:44 PM

I'm a heavy user and fable is great the #1 reason I stopped using it was the horrible safegaurd filter. I found sol close enough in capability and have only been blocked when my request was an obvious offensive cyber work. Fable blocked me on almost everything.

Optimizing a OS build? -> block

Securing a container -> block

60% is nowhere near enough for that safegaurd system. This just means I am going to be blocked half as much? Any long running task will likely get blocked.

Say you give a single big prompt and fable goes off for 6hrs of work. At hr 5 it gets blocked you now have the option of a much dumber model taking over and wrecking it or losing the entire 5hrs of work. That risk is beyond terrible and deffinetly not worth a 5-10% percieved improvement on my end. I previously would just bring sol in when that happened and realized sol is stupidly close in capability.

show 8 replies
m-schuetztoday at 5:22 AM

> Anthropic did not get much bite on Fable at its original pricing

I stopped using Fable because it kept stopping itself due to safeguards.

nsingh2yesterday at 7:05 PM

From Artificial Analysis cost per task, it looks like Fable 5.1 (max) is more expensive per task than Fable 5 (max)? Cache hit price went down, but the other components still add up to more.

Edit: 5.1-xhigh seems to be cheaper than 5-max, and 5.1-xhigh has a higher index score than 5-max. Also interesting that Fable 5.1 (high) is comparable to Opus 5 (max), but nearly half the price.

https://artificialanalysis.ai/models#cost-tabs

show 1 reply
supern0vayesterday at 7:22 PM

>Has frontier progress finally stalled?

It wouldn't surprise me if we start to see minimal performance gains from incremental changes to base models. It seems like the gains from the Opus 4.5+ incremental updates were a result of Anthropic learning a lot about post-training, the gains from RLVR, etc.

If new post-training techniques are seeing diminishing returns, we could just be back to waiting for new large pretraining runs at larger sizes for gains (even if those ultimately end up getting distilled down into smaller models because the economics for serving anything larger than Fable isn't practical).

show 1 reply
einsteinx2today at 12:17 AM

> Has frontier progress finally stalled?

From my experience using coding agents approximately 7 days per week for the past year and a half or so, we hit the top of the S curve about a year ago around Opus 4.5, and it’s mostly been harness and other tooling improvements since then with small percentage improvements coming from the actual models.

I was saying this already months before Fable dropped and thought from all the Mythos hype that maybe I was wrong…then Fable came out and was barely better than Opus 4.8.

Considering how many more parameters Fable is supposed to be than Opus, we seem to have hit a scaling limit at least with current transformer architecture considering how closely Fable and Opus benchmark and perform in practice.

rxyzyesterday at 7:32 PM

Anthropic did not get much bite because they don’t offer zero data retention with fable

anukintoday at 12:24 AM

The issue with many of these benchmarks is that it doesn’t take into the real world usage of the model. Fable for me was a step above opus. The real reason I stopped using it is because of misanthropic. I was hospitalized and asked to extend my claim to fable credits and they responded to it by denying it. I regret buying annual plan instead of monthly one.

andaiyesterday at 8:24 PM

> This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general.

Does that mean that generally available intelligence is now constrained by Moore's law? We have to wait for the actual price to come down.

andaiyesterday at 8:25 PM

> Probably leaves no room to place Opus 5.1 anywhere.

Well it'll probably be better than Fable again, lol

unsupp0rtedtoday at 12:26 AM

It's because GPT Sol is equally good and established a price ceiling

black_knightyesterday at 10:06 PM

I haven’t yet had a week without spending my Max Fable allowance. For my work (formalised mathematics) Fable is my go to for hard(ish) tasks and problems – of which I have many!

I hope they keep making it smarter! (Cheaper would be nice too, but smarter is my priority!)

Tepixyesterday at 6:57 PM

DeepSeek V4 Pro cache read pricing is $0.022 (offpeak) and

DeepSeek V4 Flash cache read pricing is $0.007

Makes it super affordable!

6thbityesterday at 7:32 PM

Huh! Yeah that feels more like an opus5.1 than a fable5.1.

arizenyesterday at 7:38 PM

Partially frontier moved to cost and speed axes.