Rounding down was my reading comprehension mistake. I should have read better.
Also, I have had extensive experience with all Gemini models except the recent few. Generally, they are always less precise and less practically useful than benchmarks suggest. You are right that here there is no evidence to think that. However, them announcing the model benchmarks without allowing access to anymore makes me think it might follow the same trend. Another point is the verbosity, I always look at that number to get the feel whether the model truly got smarter or it just wrote so many reasoning token to get there, Argon is on the higher side.
I would be very surprised if in practice, it was more practically useful than Astra, Sol, Opus or Fable.
> If they’re smart they pull a DeepSeek and teach flash to beat this entirely within a month or two, I wouldn’t bet against it.
Yes this would be a great outcome for every consumer.