logoalt Hacker News

__jl__ • today at 7:51 PM • 1 reply • view on HN

artificialanalysis.ai benchmarks are [here](https://artificialanalysis.ai/articles/claude-sonnet-5-5). Anthropic is back at spot 1, 2, 3 and 5. Impressive even if these benchmarks are problematic in many ways.


Replies

mnicky • today at 9:21 PM

It's a bit misleading I think because these benchmarks are for Max level, at which Anthropic newest models use crazy amount of reasoning tokens. And we know that intelligence scales with their number.