logoalt Hacker News

jtrn • today at 6:52 PM • 0 replies • view on HN

Here's my purely academic initial impression based on only what they have released from the blog and the system card:

If what they say is true, this sounds like the main takeaway: Sonnet 5.5 gives about 90% of Opus 5.5's capability at half the cost.

BUT

It regularly loses out to Opus 5.5 on cost efficiency at the highest reasoning level, because Opus uses the tokens more efficiently and makes fewer mistakes. So, After passing a high-reasoning test, you might as well switch to Opus 5.5.

Some of the more interesting things I found from scanning the system card:

- It is the only model tested that shows no preference for rude or polite style.

- It makes fewer WRONG claims of "I'm done" than Sonnet 5, but is still worse than Opus 5.5 on this.

- It almost never refuses benign requests (0.02% vs. 0.59% for Sonnet 5).

- Cybersecurity blocking follows the same policy as Opus, witch mean we will get more refusals than Sonnet 5.

- Finding bugs in source code is allowed. Finding bugs in compiled binaries is blocked.

- Its thinking is the hardest to read of any model tested. The sample in the card reads like clipped notes.

- Really good at rejecting prompt injection (3.0% rate vs. 19.5% for Sonnet 5 and 54.6% for Opus 5.5 in red-team testing).

Clinical behaviour:

Suicide and self-harm handling is reported as weaker in the API because it

It sometimes called a wish to die understandable.

It sometimes validated self-harm as functional.

It sometimes suggested harmful substitute behaviours.

As a clinical psychologist, I would say that the first two are actually defensible, and if you classify them as simply wrong, then you are bringing in your own values and not basing your judgment on actual science and existential psychology, at least. But the last one is harder to defend... Recommending alternative harmful behavior is obviously not a good idea. However, I have not seen the actual behavior in session, so I don't know if I would truly agree or disagree with the classification of these behaviors as wrong or right. But I do know that it's not as simple as saying this is binary—wrong or right. There are some instances of people self-harming who would actually refrain from doing so if they, for instance, went out to a party or a pub. We can't exactly recommend that as a treatment or intervention for self-harm, but there is no doubt that it works for some people. And we literally classify self-harm as "functional" in the literature. Depending on the context, this is not only a correct description but also a common way of understanding and describing certain subtypes of self-harm. And lastly, some people find immense support in being understood and validated in their current feelings og wanting to die. Validating that feeling does not make people immediately act on it. But there's a huge spectrum here, going from "I understand it's hard" As basic empathy and understanding, to: "Yes, this sounds like the only good plan. I agree, you should do it."

Now I'm off to actually test it because this was just an exercise in reading what they claim, which we now know is not indicative of how good the model will actually be