All models are benchmaxxed, period. ”Jagged frontier” is the euphemism du jour, I believe?
Anthropic/OpenAI were touting PhD-level intelligence three years ago. And they’re still shipping models that aren’t smart enough to realize things such as the need to drive the car to the car wash (because they hadn’t yet hill-climbed that particular brain-teaser).
Anthropic/OpenAI were touting PhD-level intelligence three years ago
No, they weren't. GPT-5 was where OpenAI started talking about PhD-level, and that was less than a year ago.
Jagged frontier is not the same as being benchmaxxed. Benchmaxxed is à la Goodhart's Law "when a measure becomes a target, it ceases to be a good measure." Jagged frontier is about how models that seem superhumanly intelligent at one category of tasks (e.g. coding web applications" can seem toddler level or worse at another category (spatial reasoning) because the training corpus doesn't generalize to there.