logoalt Hacker News

Gemini 4 Argon

1635 points • by bradleyg223 • yesterday at 8:04 PM • 1119 comments • view on HN

See also: Gemini 4 Argon (High): Intelligence, Performance and Price Analysis - https://news.ycombinator.com/item?id=49914236


Comments

taylorfinley • yesterday at 8:16 PM

Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.

Edit to add the fix: https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...

➕ show 23 replies
nickysielicki • yesterday at 8:19 PM

The important take away here: the leapfrogging we’ve seen this year doesn’t seem to be a temporary thing. The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back. The term he liked to use was, “concentrating”. This is yet another datapoint that he was wrong about that. AI seems more distributed amongst neoclouds and traditional hyperscalers, FAANG and startups, GPUs and ASICs than it did this time a year ago.

Nobody has a moat.

➕ show 36 replies
babelfish • yesterday at 8:06 PM

> We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.

Gemini not beating the "can't release a model" allegations

➕ show 6 replies
wg0 • yesterday at 9:14 PM

Breaking news is not the model. Breaking news is that inside Google, it is being heavily used on large code bases for writing code and it is migrating 800k lines of C++ code to Rust already.

In this space, any other company that I respect other than DeepSeek is - that would be Google. They had been honest about it from the get go including their infamous "we have no moat" memo.

This company has enormous data, their own hardware (TPUs) and their own in house experts. Actually, LLMs are invented here.

Good addition to the arsenal.

➕ show 2 replies
tazjin • yesterday at 8:10 PM

> Argon agents are working on migrating C/C++ codebases to Rust across Google

Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.

➕ show 8 replies
uvdn7 • yesterday at 8:28 PM

> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google—scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel.

To me this is way more significant than other random c++-to-rust-AI-rewrite. If they can pull it off on core C++ libraries en masse, I don't know if C++ will still be relevant in a few years.

I look forward to a post from google on this effort.

➕ show 5 replies
Revanche1367 • yesterday at 9:17 PM

Great, so they _finally_ decided to add a non-flash model and it's not available to regular subscribers for an indefinite period. What's the point of paying for the AI Ultra plan? Anthropic doing the same with Fable as far as I know, OpenAI at least allows Pro plan subscribers to use Astra. I subscribe to Gemini AI Ultra and ChatGPT Pro, and have enterprise access to Claude at work. To be fair, Gemini's flash models since at least 3.6 have been quite useful for non-complex work, but for any task where there is a bit of complexity involved, I've had to check and recheck the work multiple times myself or sometimes with another LLM to get it to follow plans accurately. It's disappointing to see yet another Gemini release ignore adding newer pro models.

Edit: seems I was wrong about Anthropic restricting Fable, I guess our enterprise plan doesn't include it. But, the block from Anthropic regarding Mythos for regular subscribers/enterprise-users is still true I think.

➕ show 4 replies
juanre • yesterday at 11:38 PM

Whatever your workflow is, make sure that model and provider are replaceable. Frontier labs will keep leapfrogging each other, as they have been doing for months.

In order for the benefits of AI to be distributed, intelligence has to become a commodity.

As long as you control the skills, the learnings, and the infrastructure setup you will be fine.

➕ show 5 replies
gopalv • yesterday at 8:10 PM

> taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.

This is good, but they're the slow mover due to this exact thing.

Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.

➕ show 4 replies
mridulmalpani • yesterday at 9:01 PM

I wonder, why Google don't make Gemini - open weights model?

Considering, Gemini 4 is in the same ballpark as SOTA models, just open source it and kill any competition from openAI and Anthropic, and be market leader.

This will be so good on so many dimensions - buying time for Google to iterate on next model, best for all folks like us, kill funding or destroy valuation of competitors and force them to be open up their model or force them to a create a much superior model than open source Gemini.

Only downside, is revenue loss from Gemini API, which I am not sure is really significant as compared to Google other revenue sources and a part of this can be captured by GCP, as you need to host the model somewhere.

➕ show 10 replies
GodelNumbering • yesterday at 9:08 PM

  Argon will launch at an introductory price [1] of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.

  [1] After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
===

So, they are basically offering opus 5.5 pricing. On AA, it scores around Sol 6.1 level (53) with avg cost per task $1.99 (https://artificialanalysis.ai/models/gemini-4-argon#cost-tab...) which is higher than Astra high ($1.73), Opus 5.5 high ($1.82), Muse Max ($1.60) and way higher than sol 6.1 Max ($0.72).

And this pricing is their 'discount pricing'. Add that to AI studio and Vertex's famously terrible caching, it is hard to see this as competitive. Google somehow is getting terrible advice on pricing (see also: the flash pricing fiasco)

But good to see more competition. I would happily take a 4 horse race (+google, +meta) than 2 horse race for US labs.

➕ show 2 replies
summerlight • today at 12:00 AM

I had access to this over few weeks, and in my impression this was the first Gemini model that I can offload complex tasks that I don't want to do myself because I have to do lots of domain specific researches, which is irrelevant to my daily works. Not 100% reliable, but its outcome is usually better than mine and the cost to verify the outcome is significantly cheaper than doing the task by myself.

The performance ceiling from the pre-training seems fairly high and they demonstrated impressive post-training improvements from Flash 3.6 -> Flash 3.8. If they can reproduce that in this model then this can be a good model for the next year. But the question is whether they can keep this up over coming years; they missed one pretraining cycle due to internal misallocation and it costed them several months of frontier competitions, and I still don't know if they addressed this structural problem.

➕ show 1 reply
arjunchint • yesterday at 8:20 PM

I dont get it, why even make this announcement, nothing's available and only one real benchmark for comparison?

Only theory is team wanted this out before perf/promo reviews to kick it over the line and then its not their problem

➕ show 9 replies
elAhmo • yesterday at 8:20 PM

> Quantum algorithmic optimization: Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.

Amazing breakthrough! So useful in day to day life, glad they put this as the first bullet of how it is making changes at Google.

➕ show 1 reply
prima-facie • today at 11:25 AM

Google is basically the incarnation of the Immortal Snail meme. OpenAI and Anthropic get millions of USD and are made immortal. But Google is the snail, constantly chasing them, slowly, day and night, relentlessly.

> Immortal Snail, also known as the Snail Assassin, refers to a hypothetical scenario in which a person is given millions of dollars and made immortal in exchange for being hunted down by a snail with a fatal touch for the rest of their existence.

➕ show 3 replies
deanc • yesterday at 8:25 PM

At this point I just think they are benchmaxxing and all talk and no action. I pay for AI plus because I wanted more storage, and when I go to gemini.google.com the most recent model I can use is 3.6-flash-lite. Two revisions have been released since then and they still can't put these things in the hands of customers. Why is it that other providers can get the models into the hands of customers right away? Google is meant to be the bigger tech company in the world.

I don't _want_ to use aistudio. The UX is confusing and I don't really know where it fits. Yet I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans).

➕ show 2 replies
skavi • yesterday at 8:27 PM

Interesting to see a mention of Fuchsia on a big Google announcement. Is the project still truly alive? Are the ambitions still as grand? Is the team as stacked as it used to be?

Also, a link to the rust root of Zircon in case anyone else was interested: https://fuchsia.googlesource.com/fuchsia/+/refs/heads/main/z...

➕ show 3 replies
iamronaldo • yesterday at 8:08 PM

Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price. Wow

➕ show 3 replies
robertbarbe • today at 7:20 AM

Google's way of presenting their model's performance is terrible compared to Anthropic or Openai. Google's benchmark results are not useful at all. If you look at their methodology, they only present the model's performance at the Max reasoning levels. https://storage.googleapis.com/deepmind-media/gemini/gemini_...

I don't know anyone who normally uses those models at these levels where Work output is excruciatingly slow.

➕ show 1 reply
rcr-anti • yesterday at 9:40 PM

According to Artificial Analysis, one metric is standing out significantly: hallucination rate. Beats frontier models by a good margin at 15%, while latest OpenAI are in the 40s-50s and Anthropic in 60s-70s (mostly). Other near frontiers are closer, Grok 4.7, GLM5.3, and Muse Spark 1.3 are all around 30%. Only other model I recall getting close was Minimax M3 at 18%.

bottlepalm • yesterday at 8:08 PM

Gemini is the model that is routinely borderline psychotic. It scares me. If we get paperclipped I won't be surprised if it's Gemini.

➕ show 12 replies
darksaints • yesterday at 8:22 PM

> Argon agents are working on migrating C/C++ codebases to Rust across Google

If anybody at google is reading this, please please pretty please prioritize or-tools. I absolutely love the project and use it all the time, but for the entire life of the project they've never had a repeatable working build system, and the whole SWIG framework is a nightmare to deal with. There's so much potential as an open source project, and a lot of external researchers would love to contribute, but the codebase is an example of everything wrong with the C++ ecosystem.

SwellJoe • yesterday at 8:09 PM

My girlfriend, you wouldn't have met her, she lives in Canada, has seen it and she thinks Gemini 4 Argon is amazing.

➕ show 9 replies
nonethewiser • yesterday at 8:30 PM

How is it even possible for every model to release benchmark results where they are #1 in 75% of categories? Like statistically, how many benchmarks would you expect there to be for this to be possible. Everyone can somehow show that they are empirically the best.

➕ show 2 replies
devinprater • today at 5:17 AM

> Gemini 4 Argon is already powering our internal workflows, with thousands of Googlers highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and writing quality.

Maybe they can push it to make the Android accessibility framework more responsive, especially when scrolling the screen with TalkBack, and catch TalkBack up with VoiceOver. Oh and add Accessibility Actions to apps like YouTube so I don't have to swipe through "video name", "video name button", "go to channel button", "more actions button", every, single, video.

But they won't because accessibility is something you have to actually prompt the model to do and who cares about a11y.

ddrcoder • today at 3:23 AM

It looks like Google's days of trailing the frontier... Argon.

➕ show 1 reply
pritambarhate • today at 12:33 PM

Has someone found out how to stop Antigravity from training on your data? I had checked 10-15 days ago on the pro plan (around $20 per month) and I couldn't find a clear option to turn off training like all other big players give.

So no matter how good their model is, it's useless to most of the regular users.

➕ show 3 replies
helsinkiandrew • yesterday at 8:19 PM

> Google Grapples With Employee Skepticism About New Gemini Model

https://www.bloomberg.com/news/articles/2026-09-30/google-gr...

➕ show 2 replies
jbl0ndie • yesterday at 11:07 PM

> Gemini 4 Argon is already powering our internal workflows, with thousands of Googlers highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and writing quality.

Like writing "What's new: this release includes stability and performance improvements" for every update to Google apps in the Android Play Store?

jjcm • yesterday at 8:10 PM

Big number results, and impressive pricing. That said it really feels like benchmarks have been hyper saturated these days. I’ll wait for hands on before getting too hyped that Google is back. It would be nice having more than just OAI / A\ in the running for SOTA top tier intelligence.

➕ show 3 replies
dom96 • yesterday at 8:17 PM

Why announce this if it’s not available yet? Why not at least announce when it will be released to the public?

None of the other AI labs do this. Really frustrating.

➕ show 1 reply
filearts • today at 1:42 PM

> Argon agents are working on migrating C/C++ codebases to Rust across Google — scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia Zircon kernel.

That seems like quite an interesting data point regardless of the quality of the model. Are the results of these migrations going to be put into production? That would be quite a shift!

➕ show 1 reply
moostii • yesterday at 10:29 PM

Excited to see Google competitive at the frontier level again. Hopefully they sort out their infrastructure and model versioning so that we can feel confident building production applications on top of their APIs. The capacity limitations I've experienced with them in the past have been deeply problematic.

AravPhi • today at 9:58 PM

they named it argon because after using it your token ar gon

AM1010101 • yesterday at 8:55 PM

Matches Astra on Artificial analysis at lower cost of $1.99 per task instead of $3.26. Still far more than GPT 6.1 sol at $0.79 for 1 point lower in intelligence.

I have found gemini models to have some of the nicest and easiest to read prose so I’m looking forward to trying this out. I hope the UI design has been preserved too

xnx • yesterday at 8:35 PM

Must've been in someone's OKR to ship in Q3.

pietz • yesterday at 9:01 PM

I know companies benchmaxx, but after what Google pulled with Gemini 3.8 Flash, I give zero f*cks about any numbers they report. No other model on Artificial Analysis dropped harder after they adjusted their weighting. Just look at their DeepSWE scores and then try to do any serious coding with the model.

Google is desperate. They haven't been performing in half a year. It's clear their researchers have been forced to integrate existing benchmarks into their training.

These numbers are meaningless. Shame on them.

ariwilson • yesterday at 8:42 PM

Damn way to undermine yourself in your own blog post Google:

"The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++."

Close but no cigar!

➕ show 1 reply
losvedir • yesterday at 9:46 PM

> expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens

Can someone help me understand this? I might have an out of date mental model of how these things work.

Fundamentally, LLMs output tokens 1 at a time, generating the next token from all the previous. And as the context window gets larger, this gets harder / slower / more expensive. So I get the idea of a maximum context window.

But I don't understand the point or meaning of an output token limit. I thought it was more a measure of price capping (since output tokens are more expensive) that a user could configure. I guess a model will keep generating tokens until it hits a "stop", so does this mean it's tuned to more aggressively produce output tokens? How does that fit into agentic loops. Are output token limits based on how long until it goes back to the user? Or does each "turn" of tool call, thought, tool call, thought, etc, get its own limit?

➕ show 2 replies
AbuAssar • today at 6:41 AM

> Argon agents are working on migrating C/C++ codebases to Rust across Google

this is big news, mass migration from C/C++ to rust with the help of AI agents will be the norm from now on!

bobkb • yesterday at 8:38 PM

IMHO Google first needs to make it easy for humans to find where to find the models and its documentation. With aistudio/model garden / Gemini enterprise etc it takes minutes to find the model.

scirob • yesterday at 8:16 PM

"Rolling out soon" don't let them hype without any release

vlenoach • today at 1:16 PM

Google still doesn't know what day it is when I ask the browser if street cleaning is in effect in NYC. The bigger gains in usability are coming at the harness level rather than model level.

➕ show 1 reply
waldrews • yesterday at 8:46 PM

Dear Google, please don't turn off your old generally available Pro-class model before your new Pro-class model is generally available (previous discussion https://news.ycombinator.com/item?id=49668196 )

vincengomes • today at 4:29 AM

Logan Kilpatrick announced [0] that training for Gemini 4 started towards the end of July. I'm surprised they could finalize a frontier level model within 2 months.

If an entire model training can be completed in 2 months there is really no moat for any company in this space now.

[0] https://i.redd.it/upsh5fekuleh1.png

➕ show 1 reply
yzydserd • yesterday at 8:37 PM

"argon" is derived from the Ancient Greek word ἀργόν meaning lazy or inactive.

thefourthchime • yesterday at 8:38 PM

I was just thinking, I bet if I refresh hacker news, a new model will come up.

huydotnet • yesterday at 10:31 PM

I recently subscribed to Claude, and was very unhappy about the usage limit of the $20 pro plan. Then I found a trick, since I got free Google AI Pro via my phone carrier, I use Opus 5.5 High for planning, and then dispatching agy to do works.

Most of the time 3.8 works fine, but it's a bit slow if compare to 3.7 Flash. If there's already a detailed plan, 3.7 can complete the task much faster. And the best thing about agy is the usage limit was very generous.

maherbeg • yesterday at 9:04 PM

Congrats to Google on this! I wonder when the labs will start requiring commits in spend. It must be gnarly to do capacity planning if users swap between models every few weeks.

hypfer • yesterday at 8:30 PM

My wish for Christmas is that Google releases the old Gemini models as open weights.

I miss you, Gemini 2.5 Pro :(

For real though. If they've become commercially uninteresting, that would be a pretty cool move.

🔗 View 50 more comments