The models are frequently getting worse at items that they aren’t being benchmarked for — and that’s happening more and more over time! Other people in other fields aren’t idiots, they are accurately perceiving the fact that these models are being hyper optimized for our industry, and are becoming less capable in other domains over time. Models of the same scale are massively worse at writing a broad variety of styles of prose than their equivalent from two years ago. (Models of increased scale are a mixed bag.)
Maybe you’re the one who needs breaking out of your cached beliefs.
So your answer is: ignore the progress, it’s not really happening, actually it’s getting worse.
That’s not a credible position, but there isn’t anything that I or anyone else can say to someone who simply doesn’t want to believe something.
Proof?
In my experience modern models are better at all tasks than models from two years ago, especially complex multi-step tasks.