logoalt Hacker News

exabrialyesterday at 7:57 PM9 repliesview on HN

Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke.

What they have done:

* Nerfed Fable, as many of noted it's useless

* Leverage Mythos as a marketing strategy, claiming its too good to release

* Removed thought traces, one of the only useful things to make sure your prompts are working correctly

* Continue tons of hype about how good they are without delivering, going to great lengths to publish how their model "hacked" its way out of a sandbox they misconfigured.

* Push a bunch of EU Overregulation onto the rest of the world with text watermarking, decreasing quality of answers

Last year, they were at least focused on making improvements. Nowadays its just a bunch of handwaving at the church of how good they are.

The only saving grace is Opus 4.6 is still available. Just sucks we haven't seen any measurable improvement, despite all of the ceremony.


Replies

DarmokTanagratoday at 5:40 AM

I can't listen to the GoT theme song without hearing "floppy wieners".

flaghackeryesterday at 8:22 PM

Text watermarking has no effect on output quality, it just works by changing the explicit source of randomness that is in practice always present in LLM output sampling. See for example https://www.seangoedecke.com/ai-text-watermarking-is-not-a-b....

show 2 replies
greenowlyesterday at 8:19 PM

Cut them a break. They are trying to IPO soon.

show 1 reply
skuetoday at 2:48 AM

> * Continue tons of hype about how good they are without delivering, going to great lengths to publish how their model "hacked" its way out of a sandbox they misconfigured.

That wasn’t Anthropic. Clearly not a well informed take.

onidjyesterday at 8:55 PM

What do you mean fable is useless?

show 1 reply
epolanskiyesterday at 9:00 PM

While I also agree that Opus 4.6, in some ways, was the last model that truly felt an assistant, all the following ones seem to have inverted the role, even a blind person can see that throwing difficult problems, and complex bugs at this model achieves more than predecessors.

I don't think there's nothing ground breaking, but sure it achieves and finds more, sooner.

llm_nerdyesterday at 9:39 PM

> Nerfed Fable, as many of noted it's useless

I certainly don't take AI advice from HN, but this is amazing.

Useless? Yes, the safeguards are ridiculous and obnoxious, though I can say that 5.1 greatly relaxes them (just doing a hardening of a project parallel with this comment, which 5.0 refused to do...so did Sol and Gemini, fwiw. The Gemini one is a laugh, because 3.1 pretending like it's a dangerous tool is simply ridiculous at this point), however Fable is extraordinarily useful.

It is, far and away, the most powerful programming model, in my experience. Like, crazily so. It absolutely annihilates Opus 4.6, which I mention given the incredibly weird reminiscing people are doing here.

And for that matter it humiliates Opus 5.0 as well. Opus 5 somehow seems like it's neck in neck in the major benchmarks, but there is simply no reality where that is true. Opus stumbles over everything that Fable just blazes through.

show 2 replies
NooneAtAll3yesterday at 9:48 PM

> as many of noted

please rephrase?

show 1 reply
jbs789yesterday at 8:09 PM

and yet we still have people saying the rate of change is increasing

my view is we had a leap over the last fe years and it's tapering off.

this is fine, but for the IPOs

show 2 replies