is it just me or is this one-upping each other every few days getting ridiculous secreting a whiff of desperation?
Some more discussion:
Gemini 3.6 Flash https://news.ycombinator.com/item?id=48993130
I keep saying this and people dont believe me, but I have b2b saas systems with actual agents running around the clock, and the performance/stability of the flash model is higher than most other models.
Meaning, its predictable with tool calls, wont spin off a million tools/do weird behavior, its reasonable. Even sonnet in a real world decision making scenario is not reliable, or will reason so long its incredibly expensive.
The benchmarks arent catching all the value, and most people have never actually ran an ai agent in a real context that matters
With all the naysayers on Gemini models I'm curious how many people actually use Gemini regularly?
For me, Gemini models are the most usable. Claude Opus and Mistral always try to turn queries into one-shot enormous commits, which just burns tokens, time and annoys me for something which is still wrong more often than not.
Gemini seems far better at listening to instructions and giving me what I actually want, on top of using far fewer tokens and wasting my time. Fable is the only model that's come close to Gemini Pro for me.
And as this is about Flash, it's exciting, I find Flash can usually get the right answer pretty quickly and without too much nonsense.
not a google fanboy by any stretch... though i've been thrilled with the flash line of models... i exclusively use it on high, and have found it to be a great fit for increasing productivity 10-fold while maintaining quality... sure it can't just go off and one-shot a bunch of work, but at the complexity level i tend to work at, neither can the frontier in a robust way that i can be confident in... sure i have to be in the loop more, but that helps keep me grounded and course-correct earlier before wasting tokens... and when you sufficiently spec out a coding/software problem, and i mean really document all of the critical nuance, it will successfully satisfy the constraints... the quality is rarely acceptable on first-pass, but it forces me to stay connected to the architecture more than i would be if using a frontier model... i've found this to be a happy middle-ground of productivity and awareness...
[flagged]
[dead]
[dead]
[flagged]
[dead]
[flagged]
[dead]
[dead]
tested the models on aistudio. despite that the knowledge cut off is march 2026 it still knows nothing about 2025!
you can check by asking "list notable world events in 2025, only list unplanned" on aistudio. or you can ask for Charlie Kirk, it also does not know. I tried it multiple time to ensure that I didn't not get routed to older models!
> but google has search
irrelevant, without deeper knowledge about cutting edge technologies or latest libraries, all of it suggestions are crap. even you ask it to search it will still use outdated keyword thus only getting outdated information.
in other word, what a disaster!
Google desperately needs to make some leadership changes within their Gemini team now that they've been surpassed by 3-5 open weight models and risk loosing frontier status all together in the near future.
At this point, I think google should consider becoming a hyper scaler for anthropic and open ai, and I predict that that is exactly what they do. The model is no longer the most valuable part of the stack.
whatever, dude. give gemma5