Gemini 4 Argon’s Google announcement lists an introductory API price of $2 per million input tokens and $10 per million output tokens. It also lists later rates of $4 and $20, with no end date for the introductory period. Cached input is 95% cheaper, and the maximum output rises from 64K to 1 million tokens.
Source: Google DeepMind. This card does not show the $2 and $10 rates. Those rates are in the announcement text below. Logan Kilpatrick’s launch post also states the introductory price in the post text, not on a separate price card.
Gemini 4 Argon price and limits
| Specification | Value | Snapshot source |
|---|---|---|
| Introductory input | $2 per 1M tokens | Google announcement, September 30, 2026 |
| Introductory output | $10 per 1M tokens | Google announcement, September 30, 2026 |
| Later input | $4 per 1M tokens | Google announcement; end date not stated |
| Later output | $20 per 1M tokens | Google announcement; end date not stated |
| Cached input | 95% off the input price | Google announcement |
| Maximum output | 1M tokens, up from 64K | Google announcement |
| Context window | 1M tokens shared across input and output | Artificial Analysis model record, checked October 5 |
The launch rates are a price snapshot, not a promise that the introductory period ends on a particular date. The release status page explains who Google has named for access; ordinary users could not call Argon in this snapshot.
How the 95% cache discount works
“95% off” means cached input costs 5% of the listed input rate. At the introductory $2 rate, that is about $0.10 per million cached input tokens. At the later $4 rate, it is about $0.20. The discount applies to cached input reads; it does not reduce output pricing.
For billing arithmetic, a hypothetical request with 100,000 newly processed input tokens and 20,000 output tokens at the introductory rates would cost $0.20 + $0.20 = $0.40 before caching effects. Reused input is charged at the discounted rate when it qualifies for caching; actual requests can have different token counts and cache behavior.
Output capacity is not context capacity
Google’s announcement changes the output ceiling from 64K to 1M tokens in one trajectory. The Artificial Analysis record also lists a 1M-token context window shared across input and output. These are different limits: output describes how much the model can generate in one response, while context describes the combined capacity available to a request.
A 1M output limit does not mean a request can use 1M input tokens plus 1M output tokens at once. It gives long-running work more room before the response is cut off, while quality, latency, and the practical token count for a task still depend on the request.
What the sticker price does not tell you
Artificial Analysis lists an estimated $1.99 cost per Intelligence Index task. That is a weighted evaluation estimate, not a flat customer fee and not a forecast for every workload. For source-separated scores, see Gemini 4 Argon benchmarks. For the metric behind the widely shared 15.1% figure, see what Argon’s hallucination rate means.
If you need a local model with no per-token API bill, download Gemma 4 and check its setup options. Gemma 4 and Gemini 4 Argon are separate model lines.
Try Gemma 4 online
~/gemma4 $ Try Gemma 4 in the browser playground before setting up a local installation.
Open Gemma 4 playground />

