AI news, with the context that matters.

Products & Services · ·

Gemini 4 Argon takes on long-running work with up to 1 million output tokens

Google has announced Gemini 4 Argon. We examine its output limit of 1 million tokens, strengths in work-related evaluations, and introductory and regular pricing. We distinguish early access for specialists from plans for general availability, and check the conditions behind Google's measurements and external evaluations.

Google has announced Gemini 4 Argon. Designed for extended development and research tasks, the model raises the output limit from 64K to 1 million tokens. The English announcement is dated September 30, 2026, and the Japanese announcement October 1. Google's announcement · Japanese announcement

The announcement does not mark the start of general access. Argon is initially being offered to selected cyber defenders in the Fairwind Program. Broader access is planned to begin with paying API customers and Google AI Ultra subscribers, but no start date has been given.

Official announcement image for Gemini 4 Argon
Image courtesy of Google, from the official Gemini 4 Argon announcement.

The questions to focus on are how longer outputs will be used and which tasks have shown results. Examining performance, cost, and availability separately makes it easier to choose what to test once the model is available to you.

Sources checked on October 2, 2026. This article is based on official materials and information published by evaluation organizations. LATENT has not tested Argon firsthand.

Contents

  1. The limit of 1 million tokens applies to output
  2. A high overall work score still leaves differences in strengths
  3. Results in the same table can come from different sources and conditions
  4. The Rust port example needs a clear comparison baseline
  5. Input and output prices both double after the introductory period
  6. Early access is limited to cyber defenders

The limit of 1 million tokens applies to output

In this announcement, 1 million tokens is the output limit. It should be distinguished from the context window, which determines how much input material a model can take in. Google says it expanded this limit to support extended reasoning and multistep work. Official announcement

More room for output does not mean the model will generate 1 million tokens every time. In actual use, the measures to check are the time and total cost required to finish your request. This limit alone says nothing about how quickly the model will answer a short question.

A high overall work score still leaves differences in strengths

On the public table from evaluation organization Vals AI, Argon leads the Vals Index at 68.90%, followed by Claude Sonnet 5.5 at 67.04% and Claude Opus 5.5 at 66.97%. The index combines results in finance, coding, law, and tax, weighted by each field's share of US GDP. It does not mean the model can automate roughly 69% of work across Japanese workplaces. Vals AI's results and methodology

The same table lists a cost per test of $15.68 for Argon and $21.34 for Sonnet 5.5. The rates used to calculate Argon's cost were $4 for input and $20 for output. These results should not be confused with calculations based on the introductory rates of $2 and $10.

On Google's model page, meanwhile, the rankings vary by evaluation. Selected results are shown below. Each row represents a different test, so the size of a score cannot be compared across rows.

Evaluation Argon GPT-6 Astra Claude Opus 5.5
DeepSWE v1.1 77.9% 74.1% 74.2%
FrontierSWE v2 55.0% 65.5% 62.3%
Terminal-bench 4.0 57.4% 58.2% 66.4%
AutomationBench 51.3% 41.4% 42.5%
LVBench 91.7% 87.5% 83.7%
CWE-bench v1 68.0% 68.0% 67.0%

Source: Google DeepMind's performance table, checked on October 2, 2026. Among the 3 models shown, Argon leads on DeepSWE, Astra on FrontierSWE, and Opus on terminal tasks. The 68.0% score on CWE-bench is a tie, not an outright lead for Argon.

Fixing code, operating a terminal, and automating an entire workflow require different capabilities. Rather than reducing them to a single claim that a model is good at coding, it is more useful to look at tests that resemble the work you intend to assign.

Results in the same table can come from different sources and conditions

According to Google's evaluation methodology, Argon was generally evaluated at its maximum thinking setting. Many metrics use pass@1, which measures the proportion of tasks completed successfully in a single attempt, but there are exceptions. Evaluation methodology PDF

Google measured Argon's DeepSWE score using an agent harness called mini-swe. Scores for other models were taken from public leaderboards and their developers' system cards. This is not a comparison in which all models were retested in the same environment.

Conditions also differ for LVBench, which evaluates understanding of long videos. Google says it supplied Argon with 1 frame per second, while using 800 frames for Astra and 600 for Opus 5.5. These differences stem from API constraints and need to be considered when interpreting the 91.7% score.

The Vals Index can also be checked on the evaluation organization's own public page. Metrics measured internally by Google, meanwhile, are reports from the developer. They should be distinguished from results independently reproduced under the same conditions.

The Rust port example needs a clear comparison baseline

As an example of internal use, Google describes improvements to the Rust version of the libgav1 video decoder. It reports that an Argon agent worked from an existing Rust port, replaced 32,000 lines of SIMD code, and made it 2.7 times as fast while preserving the same video output. Google's account of the example

The baseline is the Rust version before those improvements. Google describes it as approaching the optimized C++ version, not becoming 2.7 times as fast as the C++ version. For readers considering a code migration, the example raises a question about how much of the existing implementation's performance can be retained, as well as safety after migration.

Input and output prices both double after the introductory period

The announced API rates are in US dollars, all per 1 million tokens. The introductory prices and the prices after that period are shown separately below.

Pricing period Input Output
Introductory period $2 $10
After the introductory period $4 $20

A 95% discount on the input rate is listed for cached input. However, the announcement does not give an end date for the introductory pricing. Google's pricing details and footnotes

A simple calculation puts the output charge for 1 million billable output tokens at $10 during the introductory period and $20 afterward. Input and other costs are additional. When budgeting for a long task, you need to consider both the rate for each call and how many calls it takes to finish.

Early access is limited to cyber defenders

The Fairwind Program gives early access to advanced models to defenders in fields such as government, healthcare, and telecommunications. Argon access is limited to selected partners. Eligible participants can also use Argon through CodeMender, which helps remediate vulnerabilities. Fairwind Program

The program sets conditions for defensive and research use. Within an organization, access is restricted to cybersecurity, incident response, and penetration testing teams. It also requires individual user authentication, phishing-resistant multifactor authentication, and records of access and use.

Google identifies 4 areas of safeguards it is developing ahead of general availability: refusal of harmful requests, defenses against malicious instructions in external content, monitoring for behavior that deviates from the intended task, and stronger isolation of the execution environment. It also describes a mechanism that stops execution when monitoring detects a problem. Google DeepMind's safeguards

For general readers, the decisions available now are which tasks to try once access opens and how to evaluate them. Google has not said that an Ultra subscription alone provides immediate access. It makes sense to confirm the official start date, supported regions, usage limits, and API model ID before deciding on an actual deployment.