Google's Gemini 4 Argon leads 12 of 18 benchmarks, outputs up to 1M tokens and costs $2/$10 per million tokens at launch, but it goes to cyber defenders first, without guardrails.

Google DeepMind announced Gemini 4 Argon on September 30, 2026. It is the first model in the Gemini 4 family and the one Google calls its "new frontier model". The announcement, written by Koray Kavukcuoglu, SVP of Google DeepMind and Google's Chief AI Architect, aims it at long-horizon work: real-world software engineering, legal and finance work, and cybersecurity defense.
Two numbers stand out. First, the output limit goes from 64K tokens to 1M, so a single run can produce "hundreds of thousands of tokens in a single trajectory". Second, in Google's own comparison table, Argon leads 12 of the 18 benchmarks it disclosed against OpenAI's GPT-6 Astra and Anthropic's Claude Opus 5.5.
The catch is access. Unless you work on a security team in Google's Fairwind Program, you can't use it yet.
Argon is "rolling out to a set of trusted cyber defenders through our Fairwind Program". Fairwind is a small, trusted group of Google Cloud customers, government agencies and cybersecurity partners. Those defenders, along with Google's internal teams, get a version with the safety brakes removed: "we'll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities." Google says the model "can autonomously find, validate, and patch critical software vulnerabilities."
Everyone else waits. "Safely releasing frontier capabilities at this level requires a phased approach," the post says, adding that Google is "actively engaged in the U.S. government's voluntary process for pre-release model access." Paid API customers and Google AI Ultra subscribers come next, "as soon as possible". Google did not give a date. As of October 2, 2026, Argon is not available in the Gemini app, AI Studio or Vertex AI. The guardrail-free version is only for defenders, and Google says it is still working on guardrails before a wider release.
One early user has a result to show. Wiz runs Argon through Scan for Good, its free program for protecting critical public infrastructure. In an early demonstration, the model "uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide", a flaw Google says earlier frontier models had missed. On Wiz's internal black-box pentest benchmark, Argon also beats Gemini 3.8 Flash Cyber.

All of the figures below come from Google, as reproduced by VentureBeat. Nobody outside Google can run the model yet, so no one has checked them independently.
| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Opus 5.5 |
|---|---|---|---|
| Harvey's Legal Agent | 19.6% | 5.4% | 3.8% |
| AutomationBench | 51.3% | 41.4% | 42.5% |
| GraphWalks (256K to 1M) | 84.2% | 71.8% | 66.8% |
| Vals Finance Agent v2 | 65.4% | 53.5% | 58.6% |
| DeepSWE v1.1 | 77.9% | 74.1% | 74.2% |
| LVBench (long video) | 91.7% | 87.5% | 83.7% |
| FrontierSWE v2 | 55.0% | 65.5% | 62.3% |
| Terminal-Bench Science 0.1 | 57.6% | 68.1% | 63.3% |
| Terminal-bench 4.0 | 57.4% | 58.2% | 66.4% |
| PostTrainBench | 45.3% | 44.3% | 49.3% |
The legal score is the biggest gap. On Harvey's agent benchmark, Argon scores more than three times what GPT-6 Astra does. The losses tell a story too. Argon is ten points behind GPT-6 Astra on FrontierSWE v2 and nine behind Opus 5.5 on Terminal-bench 4.0, so for agents that live in a shell, it is not the obvious pick.
Engadget adds an outside data point: an Intelligence Index score of 60, the same as GPT-6 Astra and one point above the GPT-6.1 Sol model OpenAI released the same day. It also reports a 15% hallucination rate, the lowest among leading models.

Because defenders get Argon first, its security scores matter most right now. On CWE-bench v1, which measures how well a model fixes vulnerabilities, Argon "ties for first place with a top score of 68%". According to Engadget, the tie is with Grok 4.7 and GPT-6 Astra. Opus 5.5 scores 67%.

Prompt injection is where Argon pulls clearly ahead. On Gray Swan's indirect prompt injection test, attacks succeeded 0.7% of the time against Argon, 1.0% against Opus 5.5 and 8.5% against GPT-6 Astra. Lower is better. That is roughly one hijacked run in 140 versus one in 12. For an agent that reads untrusted web pages or tickets, the difference matters, and Gemini agents have been tricked this way before.

Google says it built Argon under its Frontier Safety Framework. That includes refusals that still allow legitimate dual-use research, adversarial training against injection, sandboxed environments and "misalignment mitigations that monitor Argon's chain-of-thought and actions".
For an introductory period, Argon costs $2 per million input tokens and $10 per million output tokens. Cached input is 95% off, which works out to $0.10. After that, the price doubles to $4 input and $20 output. Google did not say how long the introductory period lasts. These are the prices in Google's announcement as of October 2, 2026, and nobody outside Fairwind can pay them yet.
At the intro price, Argon costs one fifth as much as GPT-6 Astra ($10/$50) and half as much as Opus 5.5 ($4/$20). At the standard price, it costs the same as Opus 5.5. The bigger output limit has a price, too: one response that uses the full 1M output tokens costs $10 now and $20 later. Teams that let agents run without supervision should set a cap.
Google's own teams are the other group with access, and the post lists what they have done with it. In quantum research, Argon beat a published baseline for qubit-by-gate resource cost by 40% "in a matter of minutes". In Google's data centers, Argon agents read fleet-wide profiling data and applied memory optimizations. That freed more than 300 TiB of memory, and Google estimates total savings of 500 TiB to 1 PiB.
The most concrete example is code migration. Argon agents are moving C and C++ codebases to Rust, from tens of thousands of lines (re2, libgav1) up to the 800K+ line Fuchsia Zircon kernel. For libgav1, Argon replaced 32K lines of SIMD code in an existing Rust port. The resulting memory-safe decoder runs 2.7x faster than that port, with identical output. The Zircon migration is still in progress and under automated and manual auditing. Google has not shipped it.
For now, nothing you can deploy. If you build on the Gemini API, your models today are the same as they were on September 29. Do not plan a migration around a model with no public release date. Budget at $4/$20, not the intro price.
When it does arrive, the 1M output limit matters most for jobs that produce large artifacts in one run: code migrations, long reports, full test suites. Teams like these currently split work into 64K chunks and then stitch it back together. Legal and finance agents also have a reason to test it early, based on the Harvey and Vals scores. Teams whose agents mostly run terminal commands should keep using Opus 5.5 or GPT-6 Astra until independent tests confirm otherwise.
The bigger change is in the release order. A frontier lab gave its strongest model to cyber defenders first, without guardrails, while it went through the government's pre-release review, and gave it to paying developers second. It happened on the same day as OpenAI's DevDay and the White House Accord on Super Intelligence, which Sundar Pichai signed. Watch for Google to name a date for paid API customers and Ultra subscribers. Until then, the scores above are Google's own.