Google has introduced Gemini 4 Argon, marking the debut of the Gemini 4 series and standing as the company’s most advanced artificial intelligence model to date. Unveiled on September 30, Argon is built to manage extended, multi-step operations within software development, legal and financial sectors, and cybersecurity operations.
This release comes after Google opted to quietly cancel Gemini 3.5 Pro, an offering previously scheduled for release in June.
Gemini 4 Argon Tops 13 of 18 Benchmarks
In head-to-head testing against rival systems from OpenAI and Anthropic, Argon currently outperforms competitors across 13 out of 18 evaluation benchmarks. On the DeepSWE v1.1 benchmark, which evaluates extended software engineering capabilities, Argon achieved a 77.9% score, outpacing Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%.
When evaluated on Harvey’s Legal Agent Benchmark, Argon reached 19.6%, surpassing Astra’s 5.4%. Furthermore, it secured 68.9% on the Vals Index, an evaluation that weights financial, programming, legal, and tax tasks by their corresponding portions of the U.S. economy.
Nonetheless, GPT-6 Astra outperforms Argon on FrontierSWE v2 and OSWorld-2.0, whereas Claude Opus 5.5 takes the lead on Terminal-bench 4.0 and PostTrainBench. Both Argon and Astra achieve an identical 68% on CWE-bench v1.
Additionally, Google has expanded Argon’s maximum output capacity from 64,000 up to 1 million tokens. Introductory API rates are set at USD 2 per million input tokens alongside USD 10 per million output tokens, with cached input available at a 95% discount.
Google Already Using Argon Internally
For several weeks, thousands of employees at Google have deployed Argon for internal operations. A network of Argon agents successfully recovered over 300 TiB of memory across Google facilities, with overall projected savings projected between 500 TiB and 1 PiB.
These agents are additionally tasked with translating C and C++ code repositories into Rust, encompassing more than 800,000 lines of code within the Fuchsia Zircon kernel. Within the libgav1 video decoder, Argon substituted 32,000 lines of SIMD instructions, producing software that runs 2.7 times faster than the previous Rust implementation.
Also Read: ChatGPT vs Gemini: How Their Visual AI Capabilities Compare
Wider Release to Follow
Argon is currently being deployed to select cybersecurity professionals via Google’s Fairwind Program during its participation in the voluntary pre-release evaluation framework established by the U.S. government. Availability for commercial API users and Google AI Ultra subscribers will come next, though Google has not yet specified a launch timeline.
Chief Executive Officer Sundar Pichai stated, “We’re going to make it available as soon as we can and as safely as we can.”
Before rolling the model out more broadly, Google implemented defensive measures to mitigate risks related to cyber assaults, the proliferation of CBRN weapons, and indirect prompt injection vulnerabilities.




