logoalt Hacker News

felixriesebergyesterday at 6:19 PM66 repliesview on HN

(I work at Anthropic)

Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier.

Another point I expect not to get much attention until it all happens at once is science. People have been correctly excited about the many "sudden" breakthroughs LLMs are making in Maths, but some of the science benchmarks make me believe we'll soon see similar developments in other scientific domains. Fable 5.1 more than doubled Fable 5's Terminal-Bench-Science [1] score, which I think is meaningful.

[1] https://github.com/harbor-framework/terminal-bench-science


Replies

NL807today at 5:03 AM

Not sure if something like this is already on the table, but I would like to see Claude responses more in line with Simplified Technical English [1] by default. I find those writing styles a lot easier to read. This has been standardised as ASD-STE100 [2]. I've seen few people making SKILL.md files with that in mind, which works great, but having this by default without invoking the skill command would be better.

1. https://en.wikipedia.org/wiki/Simplified_Technical_English

2. https://asd-ste100.org/

show 3 replies
velcrovanyesterday at 6:30 PM

I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks.

I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since left ceased bothering with human languages, and it takes a rare sort of direct matrix-gazing savant to be able to try and horse-whisper them into doing or revealing anything they didn't already plan to do.

show 22 replies
DarmokTanagratoday at 5:36 AM

Thats great news for blogspam HN posts trying to pass themselves off as human prose.

belvalyesterday at 6:24 PM

As a fervent Claude Code user who made the switch to GPT 5.6 Sol over Opus 5 over hard-to-read prose this makes me happy. I love your product but the current models are very hard to work with if you need to do a lot of context switching. Brevity is key.

show 6 replies
PedroBatistayesterday at 6:31 PM

This post and comment makes me believe "science" is the new "code" for Anthropic now that the code advantage is mostly gone and lost for OpenAI, ie. they got much better and Claude become significantly worse over these months.

show 2 replies
neosatyesterday at 8:25 PM

Can you or someone else from A\ comment on whether the conversation style is coming to Opus 5 or a future 5.1 asap as well? Currently it seems the model has been made unusable by the way it 'speaks' and there is a clear solution where it can speak better but nothing has been done about the flagship model on Pro plans. I've literally had to work on Opus 4.8 which does not have this problem and speaks fine.

show 1 reply
5555watchyesterday at 8:06 PM

While I can't speak for everyone in academia, I personally don't feel comfortable in putting my research questions and outputs to a private website, before the idea is at least arxived. Especially as all the Fable/Mythos prompts are said to be human reviewed.

So I believe that, at least in the short run, we might be seeing breakthroughs in hard open problems or in low hanging problems which are not that interesting to spend time on.

I may be wrong, if some research labs have private contracted access to the models

show 3 replies
bryanlarsenyesterday at 7:11 PM

Does it fix my favorite pet peeve, the overuse of the wrong meaning of "fail closed"?

"Fail open" usually refers to a fuse that opens and kills power, meaning the system is inert and safe on failure.

"Fail closed" is the opposite -- system has power and is live.

Computer security people have appropriated the term but use it for the completely opposite meaning. When your work straddles electrical engineering and computer security the best way to avoid confusion is just to never use the term.

I can tell my Claude to never use the term, but of course now I'm seeing it everywhere in comments from other people and it drives me batty.

show 2 replies
irthomasthomasyesterday at 8:14 PM

A recent paper demonstrated how to retrieve decoded hidden reasoning traces. The authors found cases where Claude had memorized the answer but hid this fact from the visible response.

It's getting harder to trust Anthropic's models. Will Anthropic now stop hiding Claude's CoT from users? Deliver the tokens people paid for, and prove the models aren't plotting against them. After all, if the idea was to stop Chinese labs from catching up, it didn't work.

adastra22yesterday at 7:18 PM

As someone working in science, this belief confuses me. How (by what means) do you think Fable 5.1 will be able to make further progress in scientific domains? The problem with science is that there is no agentic harness. The agent can't test things. At best it can hallucinate something and ask if that hallucination "makes sense", but this doesn't work in science.

show 4 replies
razstertoday at 12:29 AM

Still not going for it. Once I learned I can train Qwen3.8 27B with my style of writing/grammar. Also more succinct. I cannot force myself to Claude or OpenAI outputs anymore. Its too much. Honestly don't think I will ever go back to paid.

show 1 reply
unshavedyakyesterday at 6:58 PM

And word on Opus 5.1 for writing style? I am on the edge of switching to OpenAI due to this horrid writing style. If Fable is better, great - but i can't even use that at work.

show 1 reply
theletterfyesterday at 7:05 PM

Docs engineer here. Nice to read about writing style: would you consider creating a writing benchmark at some point? I guess y'all are painfully aware of the load-bearing issues (pun intended).

show 1 reply
iamflimflam1today at 1:26 AM

I really hope the improvement in natural style is real.

When I’ve tried to adjust the output style is that initially it feels better - but that’s just because the new output is so refreshing to read after the horrible Claude output.

Unfortunately, after a short while you quickly realise that it’s just as vacuous as before the style change.

wouldbecouldbeyesterday at 7:39 PM

The main issue I have, which is partly connected to writing style, mainly with it dealing with our stupidity. Is that is actually thinks it knows better, and sometimes it does, but often it doesn't and then it keeps telling me I'm wrong and I have to argue with it. Opus 5 is more condescending then Fable, but it still is very tiring. Does fable 5.1 handle this better?

srousseyyesterday at 6:27 PM

Please bring to the other models, and also please only apply the AI text watermarking only to EU citizens. I may not be able to tell when Claude writes about things i don't know, but in CC it writes about my code and it is obvious.

show 3 replies
321ahTyesterday at 6:29 PM

How is it possible that all models from xAI, OpenAI, Anthropic, Qwen etc. win all benchmarks on each release?

Tomorrow all of the above (except Anthropic of course) will bump version numbers and be at the top of HN winning all benchmarks.

Science breakthroughs incoming? First of all, you are already restricting science in Fable, secondly, we have been hearing the same for several years now.

show 2 replies
fxtentacleyesterday at 8:52 PM

(I don't work at Anthropic, but I've designed RLVR tasks)

My impression is that especially for long-horizon tasks like science, the harness is much more important than people give it credit for. Claude Code + Fable 5 seems to have a tendency to "give up", get stuck in a dead end, or claim things to be impossible. But using the Fable 5 API together with a custom harness, it'll happily try 200+ variants and fail its way towards the goal.

If you give the AI a way to give up, eventually it will. If you remove that option from the harness, then thanks to the non-determinism inherent to LLMs, you get to explore pretty much all related solution attempts.

loloquwowndueotoday at 12:44 AM

What’s your honest take on how load-bearing its use of em-dashes is now? Measured, not guessed.

ryandvmyesterday at 10:06 PM

Great. I'm looking forward to it being less obvious that my colleagues have stopped understanding their jobs.

ddahlentoday at 12:45 AM

The writing style has significantly improved, however the token burn rate for tasks I have been working on seems to have skyrocketed. It definitely appears more capable (though I am unclear how much of that is just me liking the English it writes now vs actually more performant). I was using Fable 5 for some mathematical analysis assistance and redoing a part of it with 5.1 burned 60% of my session at a much faster rate.

show 2 replies
a2ff6eeb0yesterday at 7:08 PM

Nice, I'm looking forward to the improved writing on the majority of articles posted here.

MassiveOwlyesterday at 7:18 PM

Thanks! This is encouraging. I try to use Claude Code for producing client facing presentations that are static html files with charts, tables, and annotations. It never gets the tone correct and phrases things so weirdly - it drives me mad. I have to really fight it to stop it writing insights in a flowery and verbose way

bilalqyesterday at 9:05 PM

Could you share what you use internally to make Fable not sound like a word salad generator?

loftiestoday at 1:28 AM

I don't want my Claude to sound "natural". Claude is a robot and it should do behave like a robot. It should do what it's told. Nothing more and nothing less.

show 1 reply
alasanoyesterday at 6:29 PM

> It sounds a lot less stereotypically like other Claude models

Don't give me hope.

I've strained eye muscles from rolling my eyes so hard every day at how Claude writes.

Edit: first discussion with Fable 5.1 "This is the right question and it needs a real trace, not a guess."

Sigh.

show 2 replies
illusive4080today at 2:05 AM

How aware internally are employees that Opus 5’s language is incomprehensibly complicated? Please fix with Opus 5.1.

mingqiztoday at 4:22 AM

By injecting that weird prompt and not by proper post training? anthropic is truly a joke.

LtdJorgeyesterday at 7:52 PM

Please much more of that. The Claudish language makes me dizzy, and it's very difficult to steer the model to not include it.

generalizationsyesterday at 9:46 PM

> similar developments in other scientific domains

The classifier is too strict. It's rare to be able to complete a project without being permanently relegated to Opus. I'd expect that the domains where this accelerates progress will be fairly limited.

crowdyriveryesterday at 8:53 PM

Can't wait for the distillations! I'd love improvement on writing on cheap models

matheusmoreirayesterday at 10:01 PM

But is the model actually going to answer hard questions when we ask them? Or are you going to keep downgrading the models so as to avoid "uplifting" lesser lifeforms like us?

vessenesyesterday at 7:06 PM

Felix, just poking at this, and it is MUCH more pleasant to talk to, thanks to your teammates for the work.

areoformyesterday at 6:38 PM

Hey Felix,

I'm really glad for that! And I appreciate that you're making yourself available. I really do. Outreach is amazing. And thanks for making Claude.

I really do love Claude. In some ways, I'm asking this question because of just how much I am grateful for the role Claude has played in my life.

    > Fable 5.1 more than doubled Fable 5's Terminal-Bench-Science [1] score, which I think is meaningful.
But my honest question is, can I use Fable like that? Can I use Fable to do science?

To borrow a Claude-ism, this is "load-bearing" because Claude's response has been degraded for innocuous research projects concerning population-level analyses of astronaut health.

These "safety filters" trigger on questions about rabbit sex, smartphone accelerometer data to classify cat purrs, and so much more. What exactly does this score mean for users like me if it's unusable for middle school physics, biology and chemistry?

Second, I would happily quantify it for y'all, but qualitatively it feels like Fable's performance is noticeably poorer than initial release / launch.

And I am wondering if this is the case particularly for me because I use Claude via Claude Code to make a personalized care dashboard for my doctors to help me in managing my care.

As I noticed in the upgraded filter announcement, https://www.anthropic.com/news/improving-fable-5-s-biology-s...

    "In the case of Fable 5, when a classifier fires, the model re-routes the user’s request to Opus 5, a capable model that does not have the same level of biological capability as Fable 5 and which cannot provide as much assistance to a malicious user. This is the fallback that users see when their requests are blocked."
I hope that I'm off base here, but I noticed that the post avoids saying that the user is informed every time when such re-routing occurs. Would you be open to confirming whether or not this is the case?

Is the end user informed every time their query is re-routed?

Or, can you confirm that there aren't scenarios where a user's outputs are degraded without telling them? I recall that this was something that had been adopted as policy for AI research during Fable's launch.

I sincerely hope that covert response degradation is no longer practised as policy.

Sorry for putting you on the spot, but again, as Claude would say, it's because Claude's load-bearing in my life. ;)

show 1 reply
synergy20yesterday at 10:35 PM

claudism really sucks, Gemini and codex output so much better, way more like a real human being.

internet2000yesterday at 8:43 PM

Are you guys nerfing Fable 5 to make it cheaper? I know you probably can't admit to it in public, but my email is on my profile.

Waterluvianyesterday at 7:21 PM

How much of the language style outcome is a well-crafted result vs. being a somewhat unpredictable outcome of mucking with levers and knobs for a while?

azalemethyesterday at 8:12 PM

Thank you for commenting here and having the guts to face the nerderati!

I'm a Claude Max user. I've never been able to use Fable as my work in medical physics involves both particle physics, biochemistry and biology from Python bivitticus to clinical medicine. I am not a US citizen and work in Europe.

Will Fable 5.1 work on any of my problems? Fable 5 refuses outright. Is there anyone I can ask for a review or adjustment of the safeguards? It doesn't seem so, but with Opus at least I'm pretty sure I can infer lots of your training data from now precise they are. Fable is basically useless infuriatingly. I'm just finishing a proper clinical trial in ovarian cancer and trying to make a simulation environment related to our technology.

neutrinobroyesterday at 8:46 PM

Both a fable and mythos release? I'm glad to see you take the belt-and-suspenders approach seriously!

kvakkeflyyesterday at 6:22 PM

I hope not! Then my t-shirt is no longer accurate :D

Bluesteinyesterday at 6:37 PM

  ⎿  You've hit your session limit · resets 2:51am (123°24′W Etc/GMT+8)
  /upgrade to increase your usage limit.
digitaltreesyesterday at 11:20 PM

And that’s what changes the whole game — Claude

_kidlikeyesterday at 7:14 PM

Do you know if Opus 5.1 is coming and will have improvements in writing style too?

show 1 reply
jbverschooryesterday at 6:30 PM

Will it respond within a reasonable timeframe?

It’s like we’re on a 14K4 modem when there’s broadband

jesse_dot_idyesterday at 7:26 PM

I had just assumed this model would read differently due to watermarking.

Trasmattayesterday at 6:26 PM

> More work to be done (and we will!) but reading better prose makes me so much happier.

I assume this work will be done for Opus as well? Opus has seemingly gotten progressively worse at its prose and technical writing with each version. I've stopped using Claude entirely for now, because it manages to turn even the simplest technical explanation into the most obtuse and obfuscated word salad imaginable. People originally adopted Claude because it felt pleasant to use in comparison to ChatGPT, but I feel like that's really been lost (at least with the Opus line).

I feel dread when I see a wall of text generated by Opus. Every developer I've talked to feels similarly right now.

show 3 replies
yoanwaidevyesterday at 9:14 PM

as an anthropic employee, do you trust the benchmarks?

philipwhiukyesterday at 6:23 PM

It's being written with Claude so I'm wondering how much of that is just using the repo as training data: https://github.com/harbor-framework/terminal-bench-science/c...

behnamohyesterday at 6:21 PM

At this point, I don't believe a word from Anthropic employees; you guys have lost all the goodwill that you accumulated over months last year.

show 2 replies
latentseayesterday at 6:33 PM

Qwen is all you need.

show 1 reply

🔗 View 16 more replies