On a related note, I see all these quantitative benchmarks and the models getting really good at them over time. One thing I've been wondering: if the GPT series of models performs so well quantitatively, why do I still kind of hate using them relative to Claude? There’s a missing “vibes” or “taste” benchmark I think.