Gemini 4 Argon High and GPT-6 Astra Max both score 53 on the Artificial Analysis Intelligence Index snapshot from October 5, 2026. The detailed evaluations split: Argon leads AutomationBench-AA, while Astra leads Terminal-Bench 4.0, GDP.pdf, and AA-LCR. Argon also has the lower estimated cost per task. Those results suggest different task fits, but they do not prove a universal coding winner.
Source: Logan Kilpatrick, from Google’s September 30 launch charts. This AutomationBench column is Argon 51.3%, Astra 41.4%, Fable 5.1 31.4%, and Opus 5.5 42.5%. It is not Artificial Analysis AutomationBench-AA. The table below is the AA snapshot, where Argon and Astra are listed at 78% and 68%.
Argon High vs Astra Max at a glance
The Artificial Analysis comparison uses Gemini 4 Argon High and GPT-6 Astra Max. The figures below are that page's October 5, 2026 snapshot; the cost column is an estimated evaluation cost per task.
| Evaluation | Gemini 4 Argon High | GPT-6 Astra Max |
|---|---|---|
| Artificial Analysis Intelligence Index | 53 | 53 |
| Estimated cost per task | $1.99 | $3.26 |
| AutomationBench-AA | 78% | 68% |
| Terminal-Bench 4.0 | 57% | 59% |
| GDP.pdf | 22% | 31% |
| AA-LCR v1.1 | 80% | 81% |
The equal index score hides the useful detail. Argon is ahead by 10 percentage points on AutomationBench-AA. Astra is ahead by 2 points on Terminal-Bench 4.0, 9 points on GDP.pdf, and 1 point on AA-LCR v1.1.
What the scores suggest about task fit
These are inferences from the evaluations, not product specialties assigned by the vendors:
- Choose Argon for a workflow where end-to-end automation is the main test. Its 78% AutomationBench-AA result is higher in this snapshot.
- Test Astra first when terminal execution, document-grounded reasoning measured by GDP.pdf, or the AA-LCR set is central. Astra leads all three of those columns here.
- Treat a tie at 53 as a reason to inspect the task-level results, not as evidence that the models behave identically.
The numbers do not support claims that one model is inherently a frontend model or the other is inherently a backend model. Real repository work also depends on prompts, tools, context, tests, and recovery from failed actions. Run a small version of your own workload before committing to a provider.
Token rates and task cost answer different questions
The same comparison lists these token rates for the evaluated configurations:
| Price item | Gemini 4 Argon High | GPT-6 Astra Max |
|---|---|---|
| Input / output per 1M tokens | $2 / $10 | $10 / $50 |
| Cached input per 1M tokens | $0.10 | $1.00 |
The $1.99 and $3.26 figures are weighted evaluation costs, not flat fees for every customer request. A long agent run can use a different number of tokens and tools than the tasks used to calculate the index. Compare both the rate card and your measured cost per completed task.
Access still affects the practical choice
Google's September 30 announcement says Argon first goes to trusted cyber defenders through Fairwind and is used internally, with wider release beginning for paid API customers and Google AI Ultra subscribers. It does not specify an API-before-Ultra order. Check the current Argon release details before treating the benchmark comparison as a purchase decision.
Related reading
Try Gemma 4 online
~/gemma4 $ Try Gemma 4 in the browser playground before setting up a local installation.
Open Gemma 4 playground />

