There must be a deeper read on why Google can rapidly ship better small models while being delayed months on the bigger model.
What's the simplest explanation?
Is anyone here using using these models via google subscription (not api). I tried to in the past using gemini cli and then agy - headless invoked by codex and claude code, but they were so incredibly buggy that it stalled 1/2 times and I cancelled. Interested to know if that has changed!
Do they officially support you use their AI Pro subscription (or whatever the heck it's called this month, the one that gives you models in antigravity) in a 3rd party harness?
The iteration cycle is becoming very quick. Gemini 3.8 Flash arrived just 20 days after 3.7 Flash.
Similarly Qwen3.8-Max was updated in just 30 days (to the 0902 release) and Muse Spark in just 28 days (to the 1.3 release).
A year ago iterative releases were every 3-6 months. At what point will they reach nightly candidates?
I see benchmarks beating sol terra and sonnet. But is actually better? Has someone used it? I don't see actually much people that use Gemini for coding.
Is the Gemini CLI still terrible compared to Claude Code and Codex? The harness the main thing holding back Google models as they could've been the best given all the advantages in compute capacity and training data they initially had, where now even the Google CEO said they're falling behind in agentic tasks, which is sort of a vicious cycle because RLHF relies on human usage.
On my short tests: This model is amazing and the speed makes it feel like another sort of AI.
But it's bad at code reviews (maybe it's the harness agy cli?). Could not get it to same quality level on reviews like Opus, GPT 5.6, Grok. Even tried special code review skills but no luck.
Flash Cyber sounds like a villain from a 90s hacker movie and I'm here for it.
Dear Google, Kindly make you chat window on the right side of vscode in antigravity extension, There is a reason others kept it like that. I can see the code and inspect the files changed while Agents keep working. its critical for me personally.
If Google has the juice and wants to win, they need to start releasing world models.
I’m interested in a general knowledge model (closed or open weight) and not coding specific. I want to plan for travel and trip. Do you have one of your favorite HN crowd?
Hopefully before they release 4.0 Flash we will finally get Gemini 3.5 Pro.
A company with 400+B revenue from software cannot build a usable command line cli for its vital AI model?
From personal experience it feels much more capable than 3.7 Flash.
It’s strange that its score on Terminal‑Bench 4.0 is so low. They aren’t fast enough to benchmaxx that section.
I asked gemini 3.8 high to review the site I'm working on for points of high cpu/ram consumption - it failed spectacularly and also halucinated the server i/o limits
model card https://news.ycombinator.com/item?id=49537354 (doesn't 404)
Google keeps flashing everyone where everyone is expecting to get PRO'bed.
Why is it still such a bad coding agent? Does anybody have any insight?
I am continually impressed with Gemini's chat responses, which encourages me to test their agentic capabilities and... no... no... and no... every single time.
It's terrifying watching it, really.
disclaimer : Introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.
Everyone is censoring models now with anything remotely resembling cyber or bio. I already have problems with my research in mathematical epidemiology because of that - both Sol and Fable simply refuse. They keep pushing people towards Chinese models that can be decensored.
I use Gemini to make sense of things Claude says to me.
Am I the only only one thinking that Google might still "win" the AI race, despite the apparent gap?
They're apparently evolving slower than most SOTA models but "slow and steady wins the race" is probably still a thing.
And since Google doesn't depend exclusively on AI models, they can probably afford to "wait and see" where all this craze is heading.
A reminder that google is the only major lab without a meaningful opt-out of training on your data. The only way to opt out is to disable message history entirely, which seems like a darkest of dark patterns to get users to leave "opt in" to training on, because next to nobody wants to use it without message history.
shows up in /models though and encourages you to use it over 3.7 Flash I prefer this over reading specs: the "just show me" way
I wish google to thrive
Gemini 3.8 flash thinks Entoloma sinuatum is good to eat... Otherwise feels great
Nice surprise. In a few of my own tests it seems maybe a tad slower than 3.7 (but still way faster than any other LLM I've used) and even smarter. With 3.7 I felt I could just not use 3.1 Pro at all and 3.8 seems even better.
Could someone explain to me why it matters if google has the best model? Isnt the real metric cost per task?
The race to the bottom continues
What did it say?
Meh, not any noticeable improvement and unlike 3.7 high it eats all your credits, perhaps medium would be better
I have to say:
The Google brand remains powerful on HN!
I’m shocked.
It's a shame Google crams it ham-fistedly into search results and that Google has some of the reputation it has because I actually really enjoy Gemini and I don't even use it for the reason people often list which is that you can cross-reference it to stuff in your Google account
I would really love to be able to use these Gemini models in Opencode or Pi with my existing Google AI Pro subscription.
3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just increased the thinking budgets...
3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash
Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just uppsed the thinking budgets...
3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash
Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
Whatever they’re using within the Maps app is not good at all. I cannot just ask it for things conversationally like I do with ChatGPT. They really need to put a better model in there. I don’t even think it maintains context across two different queries within the same session. It’s not seamless and doesn’t just “get it” like ChatGPT does.
Yesterday I asked for food stop on my road trip 45 minutes from the current time and it gave me some options, but then I changed my mind and specifically asked for Asian restaurants and it completely forgot about the 45 minutes and gave me the closest Asian restaurant to me.
>"safety performance" - this starting to get long in the tooth. Gemini cut programming session 3 times for "safety reasons" yesterday for mentioning image generation (I need to generate bunch of those for infinite zoom virtual training app experience). After I got creative and managed to trick it to answer t was of course because "think of a children"
And in my other app I was debugging and using OpenAI to optimize some path it cut me off numerous times because it did not like JIT functionality (this is my commercial business rule evaluation engine that compiles rules to executable code inside the app to increase performance using asmjit library)
I am basically paying for them to waste my tokens and time on these 2 tasks
Who coined the phrase "cyber" for security related things lol. It's so 1999.
love these flash models
The biggest problem with Gemini is that its performance degrades the longer you use it for coding. Is it just me?
"Page not found"...
Meanwhile I pay for Pro and still don't have access to 3.7?