Google's new Gemini 4 Argon model is going to the people who fix security holes before it goes to anyone else. Google DeepMind announced it on Wednesday 30 September, US time, and said trusted cyber defenders in its Fairwind Program get it first.

It's Google's first Gemini 4 model. The Verge points out Google hadn't put out a new flagship since the Gemini 3 series in November 2025. Paying API customers and Google AI Ultra subscribers are next, but Google hasn't said when.

What Google says Argon can do

Koray Kavukcuoglu, Google DeepMind's SVP and Google's chief AI architect, made the announcement in a blog post. He wrote that Argon was "built to sustain deep reasoning across complex, long-horizon workflows." He says it's top tier at real-world software engineering, at office work like legal and finance, and at cybersecurity defence.

Photo: Chad Davis / Wikimedia Commons (CC BY 2.0)

All the big scores here are Google's own figures. Google says Argon hit 77.9% on DeepSWE v1.1, a test of long, multi-step software work, and calls that a new record. It says Argon leads the Vals Index, which weights finance, coding, legal and tax tasks by how much each sector adds to US GDP. On Zapier's AutomationBench, which checks whether a model can finish a business task from start to end, Google reports first place at 51.3%. On LVBench, a test of understanding long videos, it reports 91.7%.

The biggest change is how much the model can write in one go. Google's raising the output limit to 1 million tokens, up from 64,000. Tokens are the chunks of text a model reads and writes. That means Argon can think and write hundreds of thousands of them in a single run, and Google says that extra room helps when it's grinding through a hard problem.

It's already at work inside Google

Google says thousands of its own staff already use Argon, and it gave real examples. A team of Argon agents dug through profiling data from across Google's machines and made memory savings in its data centres. Once those changes went live they freed more than 300 TiB of memory. Google thinks the total saving will land somewhere between 500 TiB and 1 PiB.

Argon agents are also moving old C and C++ code into Rust, a language designed to stop whole families of memory bugs. Some jobs are tens of thousands of lines in core libraries like re2. The biggest is more than 800,000 lines for the Fuchsia Zircon kernel. Google says every rewrite goes through automated and manual checks, emulation testing and review before it reaches production.

One example stands out. In libgav1, Google's open-source video decoder, Argon agents took a Rust version that already existed and replaced 32,000 lines of hand-tuned SIMD code, the sort of low-level code written to squeeze out speed. They ran round after round of experiments and studied what the compiler produced until it would speed up safe Rust on its own. Google says the result is a decoder that's memory-safe, runs 2.7 times faster than the earlier Rust port and gives exactly the same video output.

Google's quantum researchers got help too. Google says Argon beat a published baseline by 40% in minutes. It did that while cutting the qubit and gate cost of a subroutine that slows down important applications.

Why defenders get it first

Google's following the same playbook it used when it opened Fairwind on 3 September with a smaller model, Gemini 3.8 Flash Cyber. We covered that in Fairwind named the defenders. SiliconANGLE reports more than 650 organisations have joined since, CrowdStrike and Palo Alto Networks among them.

Google says Argon can find, confirm and patch serious software flaws by itself. Fairwind members and Google's own teams will get the model without its cyber guardrails, so they can use everything it's got for defence. Everyone else gets a version built to turn down harmful requests.

There's already one result to point to. Google says Wiz is using Argon in its Scan for Good initiative, which finds and fixes risky weak spots in critical public infrastructure for free. In an early test, the model found a critical flaw in hospital software used around the world. It was exposing sensitive personal information, and Google says earlier top models had missed it. On CWE-bench v1, a test of fixing security flaws, Google says Argon ties for first at 68%. SiliconANGLE reports the other model in that tie is OpenAI's GPT-6 Astra, which OpenAI keeps for vetted defenders through its Daybreak program.

"Safely releasing frontier capabilities at this level requires a phased approach," Kavukcuoglu wrote. Google says it's taking part in the US government's voluntary process for giving officials early access to new models. An Agence France-Presse report carried by The Guardian noted the launch came a day after Donald Trump hosted tech leaders at the White House. Google chief Sundar Pichai and Anthropic's Dario Amodei were there, and they signed a voluntary accord pledging to police the risks of their own AI systems.

The guardrails Google is adding

Before a wider release, Google says it's still working on four things. The first is misuse. Argon's built to refuse help with cyberattacks and with chemical, biological, radiological and nuclear weapons, while still helping with genuine science that has both good and bad uses. Google says it's getting better at watching the model's internal activity to catch misuse, and that red teams inside and outside the company tested those protections.

The second is prompt injection. That's when hidden instructions in a document or web page try to take over a model. Google calls Argon its toughest model yet against these attacks and says it leads Gray Swan's Indirect Prompt Injection benchmark.

The third is misalignment, which means a model going further than its user wanted. Google says it'll watch Argon's chain of thought and actions, and stop it when it has to. It ran a similar system during training. It says it was careful not to feed what it found back into the model, so Argon wouldn't learn to hide its reasoning from the monitor. Google urged the rest of the industry to keep model reasoning out in the open.

The fourth is locking down the sandboxes, the walled-off test areas used for high-risk training and testing. Google says it'll share how it does that with partners.

What it will cost

Argon starts at an introductory price of US$2 per million input tokens and US$10 per million output tokens. Cached input, text the model has already seen, costs 95% less than the normal input rate. Once the introductory period ends, the price goes up to US$4 and US$20. SiliconANGLE notes that's what Anthropic charges for Claude Opus 5.5.

The Gemini line had a quiet patch this year. The Verge reported on 24 September that the Gemini 3.5 Pro update Pichai promised at Google I/O in May never came out, because Google put its effort into faster Flash models. Argon is what came out of that wait. Once paying developers and Ultra subscribers get their hands on it, we'll get the clearest test yet of Kavukcuoglu's claim that Google is back at the frontier.