I had access to this over few weeks, and in my impression this was the first Gemini model that I can offload complex tasks that I don't want to do myself because I have to do lots of domain specific researches, which is irrelevant to my daily works. Not 100% reliable, but its outcome is usually better than mine and the cost to verify the outcome is significantly cheaper than doing the task by myself.
The performance ceiling from the pre-training seems fairly high and they demonstrated impressive post-training improvements from Flash 3.6 -> Flash 3.8. If they can reproduce that in this model then this can be a good model for the next year. But the question is whether they can keep this up over coming years; they missed one pretraining cycle due to internal misallocation and it costed them several months of frontier competitions, and I still don't know if they addressed this structural problem.
Can it surpass (or at least maintain) 3.1 Pro when it comes to chatting?
I've been using GPT since its 3.5 release, but starting from 5, its output always contains some meaningless nonsense or provides answers that are difficult to read just to avoid hallucinations. I find Gemini to be excellent for chatting. For example, when I'm pondering a mathematical theorem, it could provides answers that inspire me, even guiding me to think about the next question.
For me, this is the only reason to maintain multiple AI subscriptions. Gemini is irreplaceable in the realm of answering questions or chatting for learning purposes; I can leave the rest to GPT.