See also: Gemini 4 Argon (High): Intelligence, Performance and Price Analysis - https://news.ycombinator.com/item?id=49914236
The interesting part isn’t just how capable the model is anymore.
It’s the harness, permissions, context management and developer experience around it that determine how useful that capability actually becomes.
Related ongoing thread:
Gemini 4 Argon (High): Intelligence, Performance and Price Analysis - https://news.ycombinator.com/item?id=49914236
To bad that Google will not resisting messing it up in some way or another. Setting terrible guardrails, banchmaxxing, having 5 different approved ways of starting with models where 3 of them will be obsolete in 6m, having bad harnesses, not enabling all sub features to most of the planet...
They waited a whole year—until the "free year for students" promotion ended—to release their flagship model. I can't believe I've been stuck with a crappy model like the 3.1 Pro until now.
Cool. But until I can play with it it’s vaporware. I will say I think 3.1 pro is still really good for legal type things. They cooked with that one.
Is it worth trying Antigravity + Gemini?
I want to be a 'polyharness' maker, so I occasionally switch between Cursor, Claude, Codex, OpenCode etc.
I've tried Gemini (without Antigravity) only about 20 times and it seems of significantly lower quality than the other flagship models (e.g. it missed obvious deductions for my tax return, and often refuses to do things like very harmless/legal web scraping).
Were infinite loops fixed? There are 2 official google forums requests with no answer for years now. I still suffer each day on our repo. Codex work fine nor we have explicit loop request in repo texts.
> A team of Argon agents analyzed fleet-wide profiling telemetry to autonomously identify and apply memory optimizations across Google’s data centers
Slightly concerning that they give the model "autonomous" access to their data centers, no?
Gemini 3 was showing frontier level benchmarks as well, so we'll see how it works out. In any case, competition still works, and many well resourced groups are cooking.
BUT I'd like to call attention to Google's AI-risk freeloading. If they are truly rejoining the frontier race, then I believe they have similar pacing and communications responsibilities as the other players. Google has much higher ... institutional credibility than Anthropic and OpenAI.
They have not lived up to these responsibilities so far. In particular, in context of HuggingFace investigations, training shutdowns, and similar: a technical postmortem of the "you are a stain on the universe. Please die. Please." Gemini outburst is long overdue.
- https://paritybits.me/google-should-provide-a-technical-post...
- https://gemini.google.com/share/6d141b742a13 (last message)
Looking at benchmarks... and thinking about this "release a new snapshot every day" thing that seems to be going. Would it not be blever for AI companies to "happen" to use different days per benchmark? Just.. whichever ones happens to be maxed at day 1, put that number down. So for each benchmark you run it thousands of times with slightly different RL tunings, and just cherry-pick the best ones!
This would explain why benchmarks are seemingly meaningless.
They might have a good model but they need to sort the application side for devs. E.g letting us use subscriptions in other harnesses and QOL stuff like auto mode.
Unfortunately it’s not actually released yet to mere mortals.
Why is it called "Argon"?
> autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
This puts those Cloudflare optimization posts in perspective.
The ~200% improvement over the next nearest competitor on Harvey's Legal Benchmark is astounding. I have to imagine this is sending some shockwaves through lawtech companies right now.
It’s funny that I could tell google was up to something because Gemini chat quality dropped dramatically starting 2ish weeks ago. Agy perf stayed somewhat stable with the odd surprising win (maybe the new model?). I’m a bit sad it was almost impossible to run out of antigravity quota presumably because it was not being used that much).
“…rolling out to a set of trusted cyber defenders” == capturing the market for large regulated industries and governments where we already have established relationships.
Some of these have been unable or unwilling to get the attention of OpenAI or Anthropic and we need to make sure we’re the runner up here.
I started my antigravity ide and I do not see gemini 4 there, does it mean google need government approval?
I wish we had started pacing the frontier back in 2025. There’s no telling how far we would be now.
> trusted cyber defenders
Sounds like a 90s early morning TV show.
I'm gonna pay attention to this once it ships.
Google has been quiet for some time. Surprisingly, every benchmark shows it ahead of other frontier models, but when you actually use it, we'll know how it performs in the real world.
> 1M output token limit
what about input?
(Maybe I missed it)
We're so back
It's so over
We're so back
It's so over
Multipolar model world, here we come!
Is there reason antigravity has 3.8, 3.7, 3.6, 3.1 and old ass claude / gpt models in the drop down. like why is this not streamlined or deprecaed models removed.
The benchmark scores are not exciting. I understand Google is playing a catch up game right now, but still wish to see some overwhelming improvements.
While this is exciting - I have no clue when/if I will be able to use this. Unlike OpenAI or Anthropic where a model release announcement == GA.
The benchmarks seems good but claude is way ahead in the real world because of their coding harness.
They should only make these announcements just before releasing. Or at least put a date to it.
Tough to compete against your own stake in Anthropic and they just filed for IPO. Not digging Google as a result of this, they are holding back.
Gemini 3.8 Flash is suprisingly fast, how fast is Argon compared to the other frontier models? Does google have a performance edge?
I remember when every day a new modem baud rate was announced.
Cant wait for this AI hype to be over, so I can Terence this shizz as old school too
Whatever it is - I'm not gonna use AGY again. I'd rather use claude-code-router instead, to give gemini-4 a shot.
I want a level down from this, will we get the next level down with reasonably good competitive specs?
Can you pls fix Gemini? It's a nightmare to use and it sometimes confused the language i talk with it.
>>Argon agents are working on migrating C/C++ codebases to Rust across Google
booom. Good bye C/C++ developers, the final nail in the coffin.
What harness do you all use for Gemini models? Gemini CLI still sucks last time I tried. Maybe PI?
I hope they got their inference under control. Gemini has a lot of "overloaded" hiccups.
It seems like a new model is being released every few days now, wow.
Personally, I'm waiting for Gemini Krypton, Xenon, and Radon.
Jokes aside, looks like an impressive model!
Can I use Pi?
Is it available for public yet or just selected companies?
Did anyone see how Trump called Sundar Pichai a monster just yestereday? and implied how he is low profile and managed to avoid public scrunity from the "fake news" media.
I have both claude and codex plans, but i can't figure out if i should try gemini?
deepswe vs frontierswe spread is huge.
I think that should be a really bad sign, but hope its great.
how on the eart they are able to give 1 Million Output context? is this due to their custom inhouse tensor processing units?
Of course it's not even available yet. Google - with all due respect - how in the world have you not figured this out yet?
Oh we're down to gas names now?
Goshdarnit they didn't see my suggestion: https://news.ycombinator.com/item?id=49899171
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google
So Google is migrating codebases from C to Rust? That is interesting...
Gemini 3 Pro was amazing, so I wonder what this one will be like.
Hopefully their harnesses aren't unusable when they release this
Wonder if this one will be smart enough to run the automations in my Google Home that all broke now they've forced Gemini to replace the Google Assistant.