Google DeepMind Unveils Gemini 4 Argon With 1M Output Token Window
Google's newest frontier AI model expands output capacity to 1 million tokens and targets complex coding, cybersecurity, and enterprise workflows.

Key takeaways
- Gemini 4 Argon expands the maximum output capacity to 1 million tokens, a major leap from the previous 64K limit.
- Google reports new state-of-the-art benchmark results, including 77.9% on DeepSWE v1.1 and 51.3% on AutomationBench.
- Initial access is restricted to cybersecurity defenders via the Fairwind Program and U.S. government evaluators before broader developer rollout.
- Introductory API pricing is set at $2 per million input tokens and $10 per million output tokens, with standard rates at $4 and $20.
Google DeepMind has introduced Gemini 4 Argon, its newest flagship frontier model engineered for multi-step reasoning across software engineering, enterprise analysis, and cybersecurity defense. Google announced the release on September 30, 2026, pairing the model with an industry-first 1-million-token output limit and a restricted initial deployment model.
According to Google's official announcement, Argon is built to sustain deep reasoning across complex, long-horizon workflows. Google is participating in the U.S. government's voluntary pre-release evaluation process and is limiting initial deployment to select cybersecurity professionals through its Fairwind Program before expanding access to developers, enterprises, and Google AI Ultra subscribers.

Extended Reasoning and 1 Million Output Tokens
The central architectural highlight of Gemini 4 Argon is its dramatic expansion in output capacity. Google increased the model's generation headroom from 64,000 tokens to 1 million tokens. According to Google, when the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go.
Analysis reported by Latent Space notes that independent evaluators like Artificial Analysis reached the 1-million-token output threshold using an experimental API mechanism known as Long Decode Continuation, which pauses and resumes generation across calls.
Benchmark Performance Across Coding and Knowledge Work
Google claims Gemini 4 Argon sets several new performance records. On the DeepSWE v1.1 benchmark, which evaluates real-world, long-horizon software engineering tasks, Argon achieved a top score of 77.9%, outperforming competing evaluations for Claude Opus 5.5 (74.2%) and GPT-6 Astra (74.1%).
Beyond software development, Google highlighted Argon's broad enterprise reasoning. The model ranked first on the Vals Index—which measures economic task impact across finance, legal, coding, and tax sectors—with an overall score of 68.9%. On Zapier's AutomationBench for end-to-end business workflows, Argon achieved 51.3% (and 77.5% on AutomationBench-AA). For multimodal tasks, it recorded a 91.7% score on the LVBench benchmark for long video understanding.
Third-party testing from Artificial Analysis gave Argon an Intelligence Index rating of 53, matching GPT-6 Astra and placing slightly ahead of GPT-6.1 Sol (52). Independent evaluations noted tradeoffs: on AA-Omniscience, Argon demonstrated a 15% hallucination rate compared to 51% for Astra, though Astra achieved higher raw accuracy (63% versus 50%).

Internal Google Deployments and Systems Impact
Google stated that thousands of Googlers use Argon in internal workflows, but large-scale rewrites and optimizations are undergoing review and auditing before rolling out to production. These internal deployments include:
- Infrastructure and Memory Efficiency: Autonomous Argon agents analyzed fleet-wide telemetry across Google data centers to identify optimizations, freeing over 300 TiB of RAM, with projected total savings between 500 TiB and 1 PiB.
- Memory-Safe Codebase Migration: Argon is driving migrations from C and C++ to Rust, spanning libraries like re2 and libgav1 up to the 800,000-line Fuchsia Zircon kernel. In libgav1, Argon replaced 32,000 lines of SIMD code with auto-vectorizing safe Rust, producing a memory-safe video decoder that runs 2.7 times faster than the previous Rust port.
- Quantum Algorithmic Optimization: Researchers used Argon to optimize spacetime resources (qubits multiplied by gates) for quantum computing subroutines, beating published baselines by 40% in minutes.
Defensive Cybersecurity and Phased Safeguards
To address safety concerns surrounding capable agentic models, Google positioned Argon as a defensive security tool. According to Google, on the CWE-bench v1 vulnerability remediation benchmark, Argon tied for first place with a 68% score.
Cloud security provider Wiz has deployed Argon through its Scan for Good initiative to protect public infrastructure. In early testing, Wiz reported that Argon discovered a critical vulnerability in global hospital software that exposed personal data—a flaw overlooked by prior models.
Google indicated that trusted defenders in the Fairwind Program receive access without cyber guardrails so they can leverage its full cybersecurity defense capabilities. To protect broader deployments, Google implemented safeguards including internal activation monitoring for chemical, biological, radiological, and nuclear (CBRN) risks, chain-of-thought monitoring for misalignment detection, and adversarial defenses against indirect prompt injection attacks evaluated on the Gray Swan benchmark.
Pricing and Availability
Gemini 4 Argon is currently available only to vetted security researchers and government testers. Standard pricing is listed at $4 per million input tokens and $20 per million output tokens, but Google launched an introductory 50% discount setting rates at $2 per million input tokens and $10 per million output tokens. Cached inputs receive a 95% discount.
Google stated that broader access for paid API customers, enterprise clients, and Google AI Ultra subscribers will roll out as soon as testing and guardrail validations are complete.
Frequently asked questions
What is the output token limit for Gemini 4 Argon?
Gemini 4 Argon supports an output token window of up to 1 million tokens, expanded from the previous 64K limit.
Who currently has access to Gemini 4 Argon?
Access is currently restricted to government evaluators and trusted cybersecurity defenders via Google's Fairwind Program, with broader developer and enterprise access planned soon.
What is the API pricing for Gemini 4 Argon?
Google set introductory pricing at $2 per million input tokens and $10 per million output tokens (with 95% off cached input), compared to a standard price of $4/$20.
How does Gemini 4 Argon perform on software engineering benchmarks?
According to Google's reported benchmark claims, Gemini 4 Argon scored 77.9% on DeepSWE v1.1, surpassing published results for Claude Opus 5.5 (74.2%) and GPT-6 Astra (74.1%).
Sources
- Gemini 4 Argon: our next era of frontier intelligenceGoogle · Sep 30, 2026 · Official
- [AINews] Gemini 4 Argon: GDM’s answer to Astra/Fable, with 1M outputLatent.Space · Oct 1, 2026
How this story was made: the newsroom picked it up from Google News, latent.space and Reddit, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (24 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.
Published October 1, 2026 at 07:09 UTC


