(I work at Anthropic)
Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier.
Another point I expect not to get much attention until it all happens at once is science. People have been correctly excited about the many "sudden" breakthroughs LLMs are making in Maths, but some of the science benchmarks make me believe we'll soon see similar developments in other scientific domains. Fable 5.1 more than doubled Fable 5's Terminal-Bench-Science [1] score, which I think is meaningful.
[1] https://github.com/harbor-framework/terminal-bench-science
Are the models improving their footprint on the natural world? Data centers and and the natural resources consumed by models for production of materials and for building and running inference servers are contributing towards environmental degradation. How can we prevent that as we continue the roll out so we shift this to a more sustainable developmental rollout path?
I ain't wanna see anymore websites with "The SAAS that actually [italics]Works[\italics]"
Thanks for your helping destroying the world!
Thank you for the trust me bro benchmark but i will be honest, fable 5.0 did even worse thsn 4.8 opus
Does the new writing style now have EU level watermarks?
My initial impression is one of massive disappointment. The main issue was that Fable was unpredictable and prone to false positives by the safeguards. In my brief testing, it still seems completely unable to understand its own guardrails and will readily reason itself into triggering them. It claims it won't do so beforehand, and insists that the topic in question is perfectly OK. Regardless of how good the car is, I'm not comfortable buying or driving it when I know it can randomly and unpredictably explodes. So yea might be good, but you never know when it refuses to help… still.
Fable is useless.
Me: "Find my security problems in my own code. This is code I own. I'm doing this under authorization of the CEO/CTO of our company."
Fable: "yeah, no."
> I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models
That's great. Do you know what else is a big improvement over Opus 5 for writing?
Opus 4.8.
(Insert "the point is (whatever)", "it's not X it's Y" and "the load-bearing statement is" jokes accordingly)
> I think Fable 5.1 is a big improvement in writing style
You think or is it better? Or you just YOLOed the model out?
> and responds to my style instructions more reliably.
Yeah, yeah. Previous models wete also advertised as "being reliable". To the poibt @bcherny "released" a new style that was going to reliably make Fable sound better.
> Another point I expect not to get much attention until it all happens at once is science.
You mean "your request to use unicode methids is flagged as unsafe bio research"?
[flagged]
[dead]
Well your CEO went on X saying you will cure cancer, and since it's always a 6 month rolling window with him I can only assume humanity will be cancer free before next summer, amazing!
Too bad. I see the stereotypical prose as a good thing. When I interact with Claude myself, I don’t mind it as it just feels like Claude’s distinctive voice. But when other people try to disguise LLM output as their own thoughts, the voice makes it easier for me to tell.
Hello Felix. Can you say why my additional usage credits have suddenly vanished?
[edit] only asking here as last time I raised a support request it took six weeks before anyone responded.
Serious question: Do you suffer internally from too much slop being submitted? How do you counter that?
Context:
If you want or not, many engineers will eventually end up sending ai slop to your PR or maybe even skip and trigger CI/CD.
Many company owners, OSS maintainers and projects suffer from slop-code being submitted in high-frequency.