Google's way of presenting their model's performance is terrible compared to Anthropic or Openai. Google's benchmark results are not useful at all. If you look at their methodology, they only present the model's performance at the Max reasoning levels. https://storage.googleapis.com/deepmind-media/gemini/gemini_...
I don't know anyone who normally uses those models at these levels where Work output is excruciatingly slow.
I almost exclusively use Max reasoning for all my work