The magnitude of improvement in unverifiable domains is small, mostly down to models doing more careful research before answering and hallucinating less. They are more thoughtful, but I expect you'd have to drop 2 major versions of Opus before you'd start to see most people really clearly be able to differentiate them.
> The magnitude of improvement in unverifiable domains is small,
What makes you say that? What is an example of a domain where the improvement is small?
I can't think of any at all. Compare something as unverifiable as "Make good music". Models now are many times better than 3 years ago.