https://artificialanalysis.ai/models/gemini-3-7-flash
The selling point for gemini continues to be speed and particularly end-to-end response time.
I've blown away by flash 3.6's speed while Opus chugs along for _hours_ on similar tasks. I've gotten into a opus designed -> gemini implemented -> opus reviewed dev cycle recently.
You can also customize Gemini Flash. It's a niche thing benefitting few, but you can tune gemini-3.7-flash in Google Vertex (now named "Agent Platform"?)
What's the typical response time for Gemini compared to other models?
Sol high is almost the same speed if you take into account drastically lower token use. Look at the artificial analysis speed vs token use. Gemini is 7x faster but 5x more tokens. And that's with Sol high being a substantially better model.
Edit: and Sol medium actually has the same AA intelligence score as Gemini 3.7, and has >7x fewer tokens, actually making it faster
It's funny that they don't mention this at all in the marketing or tech specs when it's obviously the biggest selling point by far. Without this it would be completely irrelevant.
Worth noting that OpenAI just announced that they got the full GPT 5.6 Sol model running on Cerebras at 750 tokens per second. No announcement of the pricing though...