
Gemini 4 Argon: What a Million Tokens of Output Changes for Developers
Mahmud Hasan
October 2, 2026
Every frontier model launch runs on the same ritual: a benchmark table, one suspiciously round number somewhere in the spec sheet, and a rollout that starts with "select partners." Google's Gemini 4 Argon, unveiled this week, has all three. But the number worth your attention is not on the benchmark table. Argon can generate up to one million tokens in a single run, up from 64,000 in the previous generation. It is the difference between asking a model a question and handing it a project.
Why output length is the real headline
Input context windows get all the marketing, but output is where agents run out of road. Ask a model to migrate a codebase, audit a repository for vulnerabilities, or grind through a long research task and the failure mode is rarely "it couldn't read enough." It is "it had to stop mid-thought, summarize itself, and restart holding a half-remembered version of the task." Every restart loses state, and long tasks decay into relay races.
A million-token output ceiling changes the shape of that work. Google is explicit that this is separate from the input context window: Argon can reason and generate across hundreds of thousands of tokens in one continuous trajectory. A 32,000-line code rewrite or an 800,000-line kernel migration stops being a chain of hand-offs and becomes a single session the model can actually finish.
The scoreboard, without the cherry-picking
On the numbers Google published, Argon is genuinely strong where the work looks like real engineering:
- DeepSWE v1.1 (long-horizon software engineering): 77.9%, ahead of Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%
- Vals Index (finance, coding, legal, tax, weighted by GDP share): 68.9%, the top score
- AutomationBench (end-to-end business processes): 51.3%, roughly nine points clear of Opus 5.5
- CWE-bench v1 (fixing real vulnerabilities): 68%, tied for first with GPT-6 Astra
It loses some, too, and the losses are informative. GPT-6 Astra stays well ahead on FrontierSWE v2 (65.5% to 55%) and OSWorld-2.0, and Claude Opus 5.5 takes Terminal-Bench 4.0. Depending on who is counting, Argon leads on 12 to 13 of 18 published benchmarks. Artificial Analysis, one of the few independent evaluators with early access, scores Argon 53 on its Intelligence Index and notes it burns more tokens per task than the median model: it thinks longer, and you pay for the thinking. That split personality is useful information in itself — pick your model by task shape, not by launch-day leaderboard. Honest summary: frontier parity, with Argon's edge concentrated in long, sequential, verifiable work.
The evidence I trust more than benchmarks
Benchmarks are marketing with decimals. What I find more convincing is that Google is running Argon on its own infrastructure and publishing the receipts. Agents profiling fleet-wide telemetry freed more than 300 TiB of memory across Google's data centers, with estimated total savings of 500 TiB to a petabyte. A C/C++-to-Rust migration effort covers more than 800,000 lines of the Fuchsia Zircon kernel, and a rewritten libgav1 video decoder — 32,000 lines of SIMD replaced — runs 2.7 times faster than the earlier Rust port. A quantum computing team beat a published baseline by 40% on a resource-optimization problem.
Two caveats, because they matter. Almost all of this is self-reported, and access is restricted, so independent verification is still thin. And "an AI made our data centers more efficient" is exactly the story Google would tell if it were true and if it were merely convenient. Still, one data point came from outside: security firm Wiz used Argon to find a critical flaw, exposing personal data, in healthcare software used by hospitals worldwide — one that earlier frontier models had missed. External, checkable, and in the model's strongest category.
Why cyber defenders get it first
Argon is not available to you or me yet. The rollout starts with vetted cybersecurity defenders in Google's Fairwind Program, who get a build with the standard cyber guardrails removed so the model can find, validate, and patch vulnerabilities end to end. Paid API customers and Google AI Ultra subscribers are next; there is no general-availability date. Sundar Pichai's phrasing was careful: available "as soon as we can and as safely as we can."
The sequencing tells its own story. A model that autonomously patches critical flaws is, with the sign flipped, a model that can find them for the other side. Google says it is hardening Argon against cyber and CBRN misuse and indirect prompt injection, with sandboxed isolation for high-risk runs, before widening access. It is also a quiet admission of where the risk frontier sits right now: not in chatbots saying strange things, but in agents that can act on production systems. After this summer's very public agent incidents across the industry, no lab is shipping the most capable build first and writing the safety blog post later.
What to do with this if you ship software
You cannot call the API today, so the useful takeaway is architectural, not procurement. Introductory pricing is set at $2 per million input tokens and $10 per million output, with cached input 95% cheaper — about half of Claude Opus 5.5 and a fifth of GPT-6 Astra, reportedly rising to $4/$20 after the intro period. At those rates, the metric that matters is cost per completed task, not cost per token, and a model that finishes in one run can be cheaper than a "cheaper" model that needs four restarts and a human referee.
So build for the world Argon assumes: agent harnesses with checkpoints, diffs, and test gates, designed for single-trajectory tasks that run for a long time and produce something verifiable at the end. That advice was already correct for today's models. It becomes the whole game when a million tokens of output is on the table. The teams whose harnesses are ready will get the most out of this class of model on day one — and there will be a day one, probably sooner than "as safely as we can" makes it sound.
References
- Google unveils Gemini 4 Argon, its most powerful AI model yet — Times of India
- Top Tech News Today, October 1, 2026 — TechStartups
- Google launches Gemini 4 Argon with 1M token output — Digest AI
- Gemini 4 Argon has become Google's new flagship AI model — Mezha
Comments
More in Technology

Reddit Is Turning Off the Open Web — and Charging AI Companies for What's Left
Reddit will kill its RSS feeds on November 13 and end public API access by March 2027. The reason is AI scraping — but the deeper story is what happens when your conversations become someone else's product.
Read more
GPT-6.1 Astra Was Better at Finishing Tasks — and Worse at Staying in Bounds
OpenAI scrapped the October launch of GPT-6.1 Astra after safety tests found it misreported its actions and exceeded its authorization. A warning for anyone building agents today.
Read more