logoalt Hacker News

alexjplanttoday at 4:37 PM5 repliesview on HN

I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).

I wonder what their official explanation for this behavior is.


Replies

Wowfunhappytoday at 4:55 PM

When something is new, its capabilities feel incredible. Over time, those same capabilities become mundane, and you start to notice the flaws.

(Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)

show 1 reply
espeedtoday at 5:30 PM

They did. More than once...

Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude https://www.wired.com/story/anthropic-responds-to-backlash-o...

But it's still happening: https://github.com/anthropics/claude-code/issues/81759

show 1 reply
QwenGlazer9000today at 4:48 PM

Last time they were called out, it was a regression in Claude code itself.

At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.

himata4113today at 5:10 PM

They are deploying optimizations weekly (if not daily) with various AB tests. They don't manipulate model performance, but they do actively perform tests.

bearjawstoday at 6:54 PM

You're right to push back, and one honest caveat -- they could just be lying.