logoalt Hacker News

irthomasthomasyesterday at 5:44 PM1 replyview on HN

Why does anthropic change the set of benchmarks they use with every new model release?

https://www.anthropic.com/news/claude-opus-4-7

https://www.anthropic.com/news/claude-opus-4-6


Replies

pietzyesterday at 5:55 PM

1. Benchmarks saturate 2. They select the most impressive improvments