logoalt Hacker News

felixriesebergyesterday at 6:19 PM66 repliesview on HN

(I work at Anthropic)

Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier.

Another point I expect not to get much attention until it all happens at once is science. People have been correctly excited about the many "sudden" breakthroughs LLMs are making in Maths, but some of the science benchmarks make me believe we'll soon see similar developments in other scientific domains. Fable 5.1 more than doubled Fable 5's Terminal-Bench-Science [1] score, which I think is meaningful.

[1] https://github.com/harbor-framework/terminal-bench-science


Replies

ALLTakenyesterday at 7:07 PM

Serious question: Do you suffer internally from too much slop being submitted? How do you counter that?

Context:

If you want or not, many engineers will eventually end up sending ai slop to your PR or maybe even skip and trigger CI/CD.

Many company owners, OSS maintainers and projects suffer from slop-code being submitted in high-frequency.

blondie9xyesterday at 9:58 PM

Are the models improving their footprint on the natural world? Data centers and and the natural resources consumed by models for production of materials and for building and running inference servers are contributing towards environmental degradation. How can we prevent that as we continue the roll out so we shift this to a more sustainable developmental rollout path?

m3kw9yesterday at 9:00 PM

I ain't wanna see anymore websites with "The SAAS that actually [italics]Works[\italics]"

motbus3yesterday at 7:28 PM

Thanks for your helping destroying the world!

cantalopesyesterday at 11:09 PM

Thank you for the trust me bro benchmark but i will be honest, fable 5.0 did even worse thsn 4.8 opus

hit8runyesterday at 7:17 PM

Does the new writing style now have EU level watermarks?

jtrnyesterday at 6:54 PM

My initial impression is one of massive disappointment. The main issue was that Fable was unpredictable and prone to false positives by the safeguards. In my brief testing, it still seems completely unable to understand its own guardrails and will readily reason itself into triggering them. It claims it won't do so beforehand, and insists that the topic in question is perfectly OK. Regardless of how good the car is, I'm not comfortable buying or driving it when I know it can randomly and unpredictably explodes. So yea might be good, but you never know when it refuses to help… still.

exabrialyesterday at 6:45 PM

Fable is useless.

Me: "Find my security problems in my own code. This is code I own. I'm doing this under authorization of the CEO/CTO of our company."

Fable: "yeah, no."

show 3 replies
naileryesterday at 7:16 PM

> I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models

That's great. Do you know what else is a big improvement over Opus 5 for writing?

Opus 4.8.

(Insert "the point is (whatever)", "it's not X it's Y" and "the load-bearing statement is" jokes accordingly)

troupoyesterday at 7:47 PM

> I think Fable 5.1 is a big improvement in writing style

You think or is it better? Or you just YOLOed the model out?

> and responds to my style instructions more reliably.

Yeah, yeah. Previous models wete also advertised as "being reliable". To the poibt @bcherny "released" a new style that was going to reliably make Fable sound better.

> Another point I expect not to get much attention until it all happens at once is science.

You mean "your request to use unicode methids is flagged as unsafe bio research"?

finnnkyesterday at 6:29 PM

[flagged]

marsven_422yesterday at 7:28 PM

[dead]

techpressionyesterday at 6:41 PM

Well your CEO went on X saying you will cure cancer, and since it's always a 6 month rolling window with him I can only assume humanity will be cancer free before next summer, amazing!

comexyesterday at 6:31 PM

Too bad. I see the stereotypical prose as a good thing. When I interact with Claude myself, I don’t mind it as it just feels like Claude’s distinctive voice. But when other people try to disguise LLM output as their own thoughts, the voice makes it easier for me to tell.

show 2 replies
saaaaaamyesterday at 6:27 PM

Hello Felix. Can you say why my additional usage credits have suddenly vanished?

[edit] only asking here as last time I raised a support request it took six weeks before anyone responded.