0% read

Gemini 4 Argon API Pricing and 1M Output Specs

Oct 5, 2026

Gemini 4 Argon’s Google announcement lists an introductory API price of $2 per million input tokens and $10 per million output tokens. It also lists later rates of $4 and $20, with no end date for the introductory period. Cached input is 95% cheaper, and the maximum output rises from 64K to 1 million tokens.

Blue launch card with the Gemini mark and the words Gemini 4 Argon.

Source: Google DeepMind. This card does not show the $2 and $10 rates. Those rates are in the announcement text below. Logan Kilpatrick’s launch post also states the introductory price in the post text, not on a separate price card.

Gemini 4 Argon price and limits

SpecificationValueSnapshot source
Introductory input$2 per 1M tokensGoogle announcement, September 30, 2026
Introductory output$10 per 1M tokensGoogle announcement, September 30, 2026
Later input$4 per 1M tokensGoogle announcement; end date not stated
Later output$20 per 1M tokensGoogle announcement; end date not stated
Cached input95% off the input priceGoogle announcement
Maximum output1M tokens, up from 64KGoogle announcement
Context window1M tokens shared across input and outputArtificial Analysis model record, checked October 5

The launch rates are a price snapshot, not a promise that the introductory period ends on a particular date. The release status page explains who Google has named for access; ordinary users could not call Argon in this snapshot.

How the 95% cache discount works

“95% off” means cached input costs 5% of the listed input rate. At the introductory $2 rate, that is about $0.10 per million cached input tokens. At the later $4 rate, it is about $0.20. The discount applies to cached input reads; it does not reduce output pricing.

For billing arithmetic, a hypothetical request with 100,000 newly processed input tokens and 20,000 output tokens at the introductory rates would cost $0.20 + $0.20 = $0.40 before caching effects. Reused input is charged at the discounted rate when it qualifies for caching; actual requests can have different token counts and cache behavior.

Output capacity is not context capacity

Google’s announcement changes the output ceiling from 64K to 1M tokens in one trajectory. The Artificial Analysis record also lists a 1M-token context window shared across input and output. These are different limits: output describes how much the model can generate in one response, while context describes the combined capacity available to a request.

A 1M output limit does not mean a request can use 1M input tokens plus 1M output tokens at once. It gives long-running work more room before the response is cut off, while quality, latency, and the practical token count for a task still depend on the request.

What the sticker price does not tell you

Artificial Analysis lists an estimated $1.99 cost per Intelligence Index task. That is a weighted evaluation estimate, not a flat customer fee and not a forecast for every workload. For source-separated scores, see Gemini 4 Argon benchmarks. For the metric behind the widely shared 15.1% figure, see what Argon’s hallucination rate means.

If you need a local model with no per-token API bill, download Gemma 4 and check its setup options. Gemma 4 and Gemini 4 Argon are separate model lines.

gemma4 — interact

Try Gemma 4 online

~/gemma4 $ Try Gemma 4 in the browser playground before setting up a local installation.

Open Gemma 4 playground />
Gemma 4 AI

Gemma 4 AI

Related Guides