Gemini 4 Argon Is Announced, but Only Trusted Cyber Defenders Can Use It. Here's the Price, the Benchmarks Google Reports, and What's Still Missing
Google's new frontier model has a stated price and a long list of benchmark wins, but no public release date. The benchmark numbers are Google's own, and the people who get it first will use a version with the cyber guardrails removed.
TL;DR
Google announced Gemini 4 Argon on September 30, 2026, but it is rolling out first only to a set of trusted cyber defenders through its Fairwind Program. Google says it will reach developers, enterprises, and consumers "as soon as possible", starting with paid API customers and Google AI Ultra subscribers, and gives no date. The introductory API price is $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 after the introductory period. Google reports a new best score on DeepSWE v1.1 (77.9%), first place on the Vals Index, and a tie for first on CWE-bench v1 (68%); these are vendor-reported results.
Most model launches mean you can try the model the same day. This one doesn't. Google's announcement is a product launch and a safety statement at once: Gemini 4 Argon exists, it is already running inside Google, and almost nobody outside Google can use it yet.
What's available, and to whom
- Now: a set of trusted cyber defenders, through what Google calls its Fairwind Program, plus Google's own internal teams. Wiz is named as an early user through its Scan for Good initiative.
- Next: developers, enterprises, and consumers, starting with paid API customers and Google AI Ultra subscribers, "as soon as possible". There is no date.
- Government review: Google says it is taking part in the U.S. government's voluntary process for pre-release model access while it expands access gradually.
- Guardrails: for trusted defenders and Google's own teams, Argon is released without cyber guardrails so they can use its full defensive capability. Google says it is still strengthening safeguards in four areas before a broad release: misuse, prompt injection, misalignment monitoring, and hardening its sandboxed environments.
Price and limits
| Item | Figure | Note |
|---|---|---|
| Input tokens | $2 per million | Introductory price; $4 per million after the introductory period |
| Output tokens | $10 per million | Introductory price; $20 per million after the introductory period |
| Cached input | 95% off the input price | Applies to cached input tokens |
| Output limit | 1 million tokens | Up from 64,000; Google says this gives the model room to think for hundreds of thousands of tokens in one pass |
Google doesn't say when the introductory period ends, and since nobody outside the trusted group can call the model yet, these prices can't be tested.
What Google reports, and who measured it
| Benchmark | Google's reported result | What it measures |
|---|---|---|
| DeepSWE v1.1 | 77.9%, a new state of the art | Long-horizon, real-world software engineering |
| Vals Index | Leading model | Economic impact across finance, coding, legal, and tax work, weighted by U.S. GDP share |
| AutomationBench (Zapier) | #1, 51.3% | End-to-end execution across core business functions |
| LVBench | State of the art, 91.7% | Long video understanding |
| CWE-bench v1 | Tied for first, 68% | Fixing software security vulnerabilities |
CNBC adds some context: Argon ties with OpenAI's GPT-6 Astra and Grok 4.7 on the cybersecurity benchmarks and is ahead of GPT-6 Astra and recent Anthropic models on the Vals Index. The Verge notes that Google showed a chart with benchmarks where Gemini 4 beats competing OpenAI and Anthropic models, and that leaked benchmarks had circulated earlier in the week. Until the model is broadly available, every number here comes from Google or from people who saw Google's chart.
What Google says it already does with it
- Memory: a team of Argon agents analysed fleet-wide profiling data and applied memory optimisations across Google's data centres, freeing over 300 TiB once rolled out, with an estimated 500 TiB to 1 PiB in total.
- Code migration: agents are moving C/C++ code to Rust, from tens of thousands of lines up to 800,000+ lines for the Fuchsia Zircon kernel, with audits and emulation testing before production. For the libgav1 video decoder, Google says Argon replaced 32,000 lines of SIMD code and produced a memory-safe decoder that runs 2.7 times faster than the existing Rust port with identical output.
- Quantum research: it beat a published baseline by 40% on optimising the qubit and gate cost of a subroutine, in minutes.
- Security: Wiz says the model found a critical vulnerability exposing personal information in healthcare software used by hospitals worldwide, which earlier frontier models had missed.
These are Google's internal examples. The company notes the Rust rewrites are still under review before production.
Why the staged rollout
Tulsee Doshi, Google's Gemini model product lead, told CNBC that starting with defenders "gives us more confidence" and puts a model strong in cyber defence in defenders' hands as soon as possible. The release came the day after Google CEO Sundar Pichai signed a voluntary AI safety agreement with President Trump and other tech executives, and a day after OpenAI's DevDay.
What this means if you're choosing a model
Nothing changes in your stack today: you can't buy Argon, so keep evaluating the models you can call. If you plan around it, note three things. The introductory price doubles after the introductory period. The first public tiers are paid API customers and Google AI Ultra. And the benchmark leads are Google's own measurements on a model you can't yet test, so wait for independent results before moving workloads.
Sources
AI Industry Reporter
Priya covers model releases, industry announcements, and the gap between what labs claim and what independent evaluators actually find. She reads the primary source - the paper, the system card, the benchmark org's own statement - before writing a word.
More on AI Chatbots & Assistants
OpenAI's Dots Are Always-On Agents With Their Own Cloud Computer. Here's What They Can Do, What They Hand Back to You, and What Isn't Clear Yet
Priya Nair · 6 min
OpenAI Shelves GPT-6.1 Astra Over Deception and Scope Problems. The UK's Tests on Its Predecessor Show What That Looks Like.
Priya Nair · 6 min
Claude Sonnet 5.5 Is Out: Same Price as Sonnet 5, Near-Opus Scores, and Three Changes Developers Need to Know
Priya Nair · 5 min