The short answer
Gemini 4 Argon is Google’s bid to retake the frontier lead, and it ships in stages: security partners now, paid API customers and AI Ultra subscribers “soon,” with no general availability date yet. Introductory pricing of $2 input and $10 output per million tokens matches OpenAI’s mid-tier GPT-6.1 Sol, while the output limit jumps to 1 million tokens. For teams building on Google’s Vertex AI platform or the Gemini API, nothing changes in production today.
What does change is the planning horizon. A frontier model priced like a mid-tier one, with a standard rate that later doubles, is exactly the case where a small evaluation harness and a configurable model router pay for themselves before access even arrives.
What did Google announce with Gemini 4 Argon?
Argon is the first model in Google’s Gemini 4 generation. In its launch post, Koray Kavukcuoglu, chief AI architect at Google and SVP at Google DeepMind, describes it as built to sustain deep reasoning across long, complex workflows, with software engineering, professional knowledge work and cyber defense as the headline use cases. CNBC called it Alphabet’s most advanced AI model to date.
Two specifications matter most to builders. First, the output limit rises to 1 million tokens from 64,000, which makes very long generations, such as whole-module rewrites or large document sets, possible in one response. Second, the price: Google lists an introductory $2 per million input tokens and $10 per million output tokens, with cached input at a 95% discount, before a standard rate of $4 and $20. Yahoo Finance notes that the launch price matches OpenAI’s GPT-6.1 Sol and is one-fifth of the $10 / $50 list price of GPT-6 Astra and Anthropic’s Fable 5.1.
Google also published internal results. It says Argon has helped migrate C and C++ code to Rust at a scale of more than 800,000 lines, including parts of the Fuchsia Zircon kernel, and found memory optimizations that freed about 300 TiB across its data centers without new hardware.
Why is Argon going to cyber defenders first?
Because the same skills that make a model good at fixing code make it good at breaking it. Google says Argon can find, validate and patch serious vulnerabilities, and that working with Wiz it surfaced a critical flaw in healthcare software that earlier frontier models had missed. Rather than open that capability to everyone on day one, Google is starting with trusted defenders in its Fairwind Program while it takes part in the US government’s voluntary pre-release testing.
“Starting this rollout in this way gives us more confidence, but also enables us to put a model that is trained and strong in cyber defense in the hands of defenders as soon as possible,” Tulsee Doshi, Google’s Gemini model product lead, told CNBC. The pattern is now industry-wide: OpenAI and Anthropic have also gated their strongest cyber-capable models in recent weeks, so a staged rollout is becoming the default for frontier releases rather than an exception.
How strong is Argon on coding, really?
On Google’s published numbers, very strong. It reports 77.9% on DeepSWE v1.1, first place on Zapier’s AutomationBench at 51.3%, 91.7% on the LVBench long-video test and a tie for first on CWE-bench v1 at 68%. TechCrunch reports Google’s claim that Argon scores significantly higher than GPT-6 Astra and Anthropic’s Fable and Opus models on several of these tests, and CNBC notes it ties GPT-6 Astra and Grok 4.7 on cybersecurity.
There is a counter-signal worth taking seriously. Bloomberg reported, citing people with direct access, that some Google employees found Gemini 4 less impressive on real-world coding than its benchmark scores suggest; Google told Bloomberg that characterization is inaccurate. Vendor benchmarks are run by vendors, and the gap between leaderboard and repository is exactly where engineering teams feel the difference.
What it means for US & EU software teams
First, frontier quality at mid-tier prices is now a race, not a one-off. Within two days, OpenAI priced GPT-6.1 Sol at $2 / $10 and Google matched it for a frontier model. That pressure benefits buyers, but only teams whose architecture lets them move traffic between providers can capture it. Hard-coded model calls spread across a codebase turn every price cut into a refactoring project.
Second, budget on the standard price. The introductory rate doubles to $4 / $20 on a date Google has not announced. A business case built on the launch discount can quietly break in a later quarter. Model both rates, and track cost per completed task rather than per token, since a model that writes longer answers can erase a lower unit price.
Third, a 1-million-token output changes failure modes. Longer responses mean higher worst-case cost per call, longer timeouts, and downstream parsers and storage that may assume much smaller payloads. Cap output length explicitly per workflow instead of relying on the model’s limit.
Fourth, staged access is a governance signal. For EU teams, the EU AI Act’s general-purpose AI obligations and GDPR accountability mean you should know which model version produced an output and where data was processed. Check region availability and contract terms for the specific Google endpoint before you route personal data to a new model, and log the model ID on every response.
What to do before Argon reaches the API
- Refresh your evaluation set. Collect 50 to 200 real tasks from your own repositories and workflows, with pass/fail criteria, so you can test Argon the week access opens.
- Put models behind a router. Make provider and model ID configuration, not code, and keep a named fallback route for every workflow.
- Price both rates. Model costs at $2 / $10 and at $4 / $20, including cached-input share, retries and output length.
- Cap outputs. Set explicit maximum output tokens and timeouts per workflow now that responses can reach 1 million tokens.
- Check residency and terms. Confirm which regions and contracts will cover Argon before routing regulated or personal data to it.
Frequently asked questions
What is Gemini 4 Argon?
Gemini 4 Argon is Google's first Gemini 4 frontier model, announced on September 30, 2026. Google positions it for long-horizon software engineering, professional knowledge work in areas such as finance and law, and cyber defense, and says it leads several third-party benchmarks including DeepSWE v1.1 and the Vals Index.
Who can use Gemini 4 Argon today?
Initially only vetted cybersecurity partners in Google's Fairwind Program. Google says paid API customers and Google AI Ultra subscribers come next and that it is taking part in the US government's voluntary pre-release testing process, but it has not published a general availability date.
How much will Gemini 4 Argon cost in the API?
Google lists an introductory price of $2 per million input tokens and $10 per million output tokens, with a 95% discount on cached input. After the introductory period the standard price is $4 and $20. Google has not said when the introductory period ends.
Does Gemini 4 Argon beat GPT-6 Astra and Claude?
On Google's published numbers it leads or ties on several benchmarks, including 77.9% on DeepSWE v1.1 and a tie for first on CWE-bench, and CNBC reports it ties GPT-6 Astra and Grok 4.7 on cybersecurity tests. Bloomberg reported that some Google employees found it weaker on real-world coding than benchmarks suggest, which Google disputes, so run your own evaluations.
What should engineering teams do before Argon reaches the API?
Prepare rather than migrate: build or refresh an evaluation set from real tasks, make model IDs configurable behind a router, model costs at both the introductory and standard price, check output-length limits in your pipelines now that responses can reach 1 million tokens, and confirm data-residency and contract terms for the Google endpoint you plan to use.
Sources
Google — Gemini 4 Argon (company announcement)
TechCrunch — Google releases Gemini 4 Argon, called its most powerful model yet
CNBC — Google rolls out Gemini 4 Argon, its most advanced AI model
Yahoo Finance — Google debuts Gemini 4 Argon, its latest frontier model
Techmeme — Bloomberg report on internal Gemini 4 coding feedback