# Gemini 4 Argon Scored 53 Before Google Opened the API

**Summary:** Gemini 4 Argon scored 53 in an independent evaluation, but broad access has not started. I would design the workload test now and leave any migration decision until the paid API opens.

- Canonical: https://markhuang.ai/news/gemini-4-argon-score-before-api-opens
- Language: en
- Author: [Mark Huang](https://markhuang.ai/about)
- Published: 2026-09-30
- Section: News
- Tags: Gemini 4 Argon, Google DeepMind, AI Benchmarks, Model Evaluation, AI Pricing
- Source: [Artificial Analysis](https://artificialanalysis.ai/models/gemini-4-argon)
- License: https://creativecommons.org/licenses/by-nc/4.0/

---

![A luminous crystalline computing core sits behind glass and a closed security gate in a dark evaluation lab](https://cdn.markhuang.ai/news/gemini-4-argon-score-before-api-opens/hero.webp)

*Gemini 4 Argon has reached the scoreboard. Most buyers are still outside the gate.*

[Artificial Analysis has measured Gemini 4 Argon](https://artificialanalysis.ai/models/gemini-4-argon), and the first independent numbers are strong. The high-reasoning configuration scored 53 on its Intelligence Index. The median among comparable models on that page is 26. Google's launch, however, starts with trusted cyber defenders rather than an open API.

That mismatch changed my read on the release. Argon has moved beyond rumor and Google's own benchmark table. I still cannot put it through a purchasing test. It belongs on my shortlist today, but I would leave existing workloads and budgets alone until customers can buy the service and test it.

The rate card and the capability claims already invite side-by-side comparisons. A buyer still needs to see the available capacity, stable limits, cost on a real task, and failures inside the production harness.

## A scorecard before a storefront

The submitted source does more than repeat Google's launch chart. Artificial Analysis says its Intelligence Index v4.3.2 combines ten evaluations covering knowledge work, automation, coding, science, and long-context tasks. Its Argon run produced 110 million output tokens across the index, above the displayed median of 82 million. The weighted average cost was $1.99 per task.

The score earns Argon test time. The output total makes me curious about the route it took. Because 110 million covers the entire evaluation, it cannot forecast one job. It does give me a reason to measure verbosity, latency, and review effort beside task success.

Google's own benchmark table is impressive but mixed. The [Gemini 4 Argon model page](https://deepmind.google/models/gemini/) reports 77.9% on DeepSWE v1.1, ahead of the listed competitors, while Argon trails GPT-6 Astra and Claude Opus 5.5 on FrontierSWE v2. It also trails Opus 5.5 on Terminal-bench 4.0. That looks like a capable general model with specific strengths, not a universal coding winner.

## The price card has two clocks

[Google says Argon will launch](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/) at an introductory $2 per million input tokens and $10 per million output tokens. After the introductory period, those rates rise to $4 and $20. Google has not attached an end date to that introductory period on the announcement page I inspected.

The access clock is separate. Argon is initially rolling out through Google's Fairwind Program to a set of trusted cyber defenders. Google says developers, enterprises, and consumers will follow, beginning with paid API customers and Google AI Ultra subscribers, but the post does not give a public API date.

The advertised rate is concrete enough for a spreadsheet and incomplete as a purchasing quote. I can model both token prices. I cannot measure my workload's token use, repair rate, or capacity needs until the service opens. I made the same distinction in my [analysis of token sticker prices](https://markhuang.ai/news/token-sticker-price-is-a-trap): the useful unit is a result that passes review.

> **Info:**
>
> I would budget an Argon pilot at the announced permanent rate of $4 input and $20 output per million tokens. If the pilot only works at the introductory price, the economics are already fragile.

## One million output tokens is capacity, not proof

Google is also raising Argon's maximum output from 64,000 tokens to one million. That gives a long coding or research trajectory a remarkable amount of room. Few useful tasks should need all of it.

More output can preserve continuity across a large migration or a long investigation. It can also leave more code to review and more claims to verify. A drifting trajectory gets expensive. Artificial Analysis's finding of above-median aggregate output use does not prove waste. It makes efficiency a question worth carrying into the pilot.

Public discussion keeps returning to that gap. In one [Gemini community thread](https://www.reddit.com/r/GeminiAI/comments/1wufgo3/gemini_4_argon_our_next_era_of_frontier/), commenters praised the benchmark table, asked when they could test the model, and wondered whether the scores would survive normal use. Axios also [described the initial rollout as limited to a small group of cybersecurity partners](https://www.axios.com/2026/09/30/google-gemini-4). The benchmarks remain useful. They just cannot settle the buying decision yet.

## The launch earns a test slot

Teams do need lead time. A model with a 53 independent score, strong published enterprise results, and a million-token output ceiling may change how a company plans a difficult migration or research workflow. Waiting until general availability to think about the test would waste that time.

I would prepare the test now and leave the migration alone. First I would choose representative tasks with known acceptance criteria, record the current model's baseline, and decide which failures require human intervention. Once Argon reaches the paid API, I would measure accepted-task rate, total token use, wall-clock time, retries, and correction time. Cyber or tool-using work also needs strict permissions and a recovery path.

That follows the rule I use for the [Artificial Analysis Agentic Index](https://markhuang.ai/news/agentic-index-needs-your-failure-test): a leaderboard can choose where I spend evaluation effort, but only a local failure test can award production credentials.

Gemini 4 Argon has done enough to become a serious candidate. Its score, rate card, and giant output limit arrived before the public product. That is enough information to design a good pilot. The evidence for switching will have to come from running it.
