The benchmark also doesn't include speed. You almost think something has gone wrong when using it because it returns full responses so incredibly fast.
This one uses that as a priority weight: https://winstonrc.github.io/ai-coding-agents-leaderboard/
Not just speed, also reliability. IME, Gemini's speed and quality doesn't degrade badly during weekday working hours compared to OAI, and especially Anthropic.