Funnily enough, I just ran a task on AI Studio with 3.6 yesterday and got 3.7 to do a similar one today; so it serves as an interesting and quick comparisons between the old and the new (usually, if enough time passes between your use of one model and the next, you'll have a sourer view of it than its actual competence suggests).
It hallucinated in both cases despite being given an API key and building a lot of pipes to access data using this. It was a simple "oh shit" fix moment for the model, but weird how eager it was to hallucinate despite the process being designed for it to be data-driven.
We should move past the idea that benchmarks alone tell us whether a model is getting better. I would've had the same experience a year or two ago with 1.5, and the solution would've been similar (keep prompting). I've been investing time into making system prompts and input prompts more meticulous, but the fundamental "it will make shit up" problem still remains, even though it shouldn't when the job involves calling tools.
I know this sounds like I'm expecting superpowers of it (I'm not), but my point is just that these incremental benchmark gains may not reflect user experience.
For me, it just shits the bed: https://news.ycombinator.com/item?id=49292924
Reading the google blog and these discussions makes me feel like I'm taking crazy pills, seriously. Side-by-side comparisons with the exact same inputs or it didn't happen, that's my rule going forward. Test all the things, believe nothing.