Oct 8 edition/Reporting & analysis
ModelsCodingSafetyBusiness

ModelsArchitectures & capability

Gemini 4 Argon ties GPT-6 Astra in independent testing and hallucinates far less, but access stays limited

Google DeepMind's Gemini 4 Argon ties OpenAI's GPT-6 Astra on Artificial Analysis's independent index and hallucinates far less. It answers fewer factual questions correctly, uses more tokens, lags on terminal-agent tasks and remains restricted to a small vetted group.

THE CORE IDEAS4 TAKEAWAYS
01

Artificial Analysis tested Argon independently at its highest reasoning setting. It scored 53 on the firm's Intelligence Index, the same as GPT-6 Astra. On factual questions Argon was right less often (50% versus 63%) but hallucinated much less (15% versus 51%). That makes it a better fit where a confident wrong answer is costly than where broad recall matters most. [4] [8]

02

Google says Argon leads outright on 12 of its 18 published benchmarks and ties on one more, including 77.9% on DeepSWE v1.1. These numbers come from Google itself. Argon trails rival models on Terminal-Bench 4.0, though sources disagree on the exact scores. Bloomberg reported that some Google staff doubt its real-world coding, which Google disputes. [7] [5] [3]

03

Total cost per task depends on how many tokens a model uses, not just its price per token. Argon uses about 62K output tokens per index task, compared with 27K for Astra. At introductory pricing that makes Argon cheaper per task. Once standard pricing starts, it costs about 1.2 times as much as Astra. [4] [8]

04

Access starts with vetted cyber defenders in Google's Fairwind Program and a voluntary US government pre-release process. Paid API customers and AI Ultra subscribers come next, with no date given. Google links the staged rollout to gathering feedback and strengthening safeguards against cyber misuse. [5] [7] [6]

WHY IT MATTERS

independent testing shows Argon matching GPT-6 Astra overall with a much lower hallucination rate.

Read the full assessment

Implication: it may suit fact-sensitive work, but weaker terminal-agent results, about double the token use and no general-availability date argue for testing before migrating.

Sources are gathered and checked against each other. Below is the dossier.

Executive brief

The model leading Google's own benchmarks is also the one least likely to make things up. In Artificial Analysis's independent testing, Gemini 4 Argon ties OpenAI's GPT-6 Astra at 53 on the Intelligence Index. It answers fewer factual questions correctly (50% vs 63%) but has a much lower hallucination rate (15% vs 51%) (Artificial Analysis). Google says Argon leads 12 of the 18 benchmarks it published. It trails on terminal-agent work, Bloomberg reports some Google staff doubt its real-world coding, and most customers cannot use it yet.

What changed and event timeline

  1. Google DeepMind announces Gemini 4 Argon

    Pitched as a frontier model for long coding, enterprise and cyber-defense jobs. Access starts with vetted cyber defenders in the Fairwind Program (;).

  2. Google's benchmark table appears

    Argon leads outright on 12 of 18 benchmarks and ties on one. It scores 77.9% on DeepSWE v1.1 but 57.4% on Terminal-Bench 4.0 ().

  3. Artificial Analysis publishes its own evaluation

    Argon (high reasoning) ties GPT-6 Astra (max) at 53. The firm says this puts Google back among the top three labs ().

  4. Bloomberg reports internal doubts

    Some Google staff question how well Argon codes in practice. Google says calling it an underperformer would be inaccurate ().

  5. Coverage shifts to the trade-off between accuracy and hallucination

    TNW reports the lower hallucination rate alongside Argon's heavier token use ().

  6. Commentary repackages the launch

    A Reddit post and a YouTube video narrated by Julian Goldie's "digital avatar" present a "3 wins, 3 losses" scorecard, mixed with course promotion ().

Capabilities and access

  • Model: Gemini 4 Argon. Artificial Analysis tested it at "high," the top reasoning setting (Artificial Analysis).
  • Output: up to 1M tokens per response, up from 64K (TNW).
  • Access order: Fairwind defenders and a voluntary US government pre-release process first.
Read the full section
  • Model: Gemini 4 Argon. Artificial Analysis tested it at "high," the top reasoning setting (Artificial Analysis).
  • Output: up to 1M tokens per response, up from 64K (TNW).
  • Access order: Fairwind defenders and a voluntary US government pre-release process first. Paid API customers and Google AI Ultra subscribers come next, with no date given (VentureBeat).
  • Price: an introductory $2 per million input tokens and $10 per million output tokens, rising to $4/$20 later. Cached input gets a 95% discount (Artificial Analysis).

Technical analysis

  • Token use: Argon spends about 62K output tokens per Index task, versus 27K for Astra.
  • Undisclosed details: the reviewed sources do not describe Argon's architecture or training. A third-party check found no Argon entry in DeepMind's model-card index (OrcaRouter).
  • Reproducibility: with access restricted, nobody outside the program can rerun Google's 18-benchmark table yet.
Read the full section
  • Token use: Argon spends about 62K output tokens per Index task, versus 27K for Astra. At the discounted price a task costs $1.99 (Astra: $3.26). At standard pricing it costs $3.98, roughly 1.2× Astra (Artificial Analysis; TNW). Budget for verbose reasoning, not just the per-token list price.
  • Undisclosed details: the reviewed sources do not describe Argon's architecture or training. A third-party check found no Argon entry in DeepMind's model-card index (OrcaRouter).
  • Reproducibility: with access restricted, nobody outside the program can rerun Google's 18-benchmark table yet.

Claims and evidence

  • Leads 12 of 18 benchmarks; 91.7% LVBench (long-video understanding); 77.9% DeepSWE v1.1: vendor-reported (VentureBeat).
  • 0.7% attack success on Gray Swan's indirect prompt-injection benchmark: vendor-reported (same source).
  • Index score of 53, Omniscience results, cost per task: independent (Artificial Analysis).
Read the full section
  • Leads 12 of 18 benchmarks; 91.7% LVBench (long-video understanding); 77.9% DeepSWE v1.1: vendor-reported (VentureBeat).
  • 0.7% attack success on Gray Swan's indirect prompt-injection benchmark: vendor-reported (same source).
  • Index score of 53, Omniscience results, cost per task: independent (Artificial Analysis).
  • Wiz used it to find a healthcare vulnerability that earlier models missed: a partner anecdote (TNW).
  • The video repeats these figures (1:27, 1:59). It adds no independent testing.

Context and prior work

  • Argon is Google's first model above Flash size in more than seven months, after the company focused on Gemini 3.8 Flash.
  • On Artificial Analysis's index, Gemini 3.8 Flash scored 41 and Gemini 3.1 Pro Preview scored 30 (Artificial Analysis).
  • Main rivals are GPT-6 Astra, GPT-6.1 Sol and Claude Opus/Sonnet 5.5.
Read the full section
  • Argon is Google's first model above Flash size in more than seven months, after the company focused on Gemini 3.8 Flash. That claim comes from the video (1:09).
  • On Artificial Analysis's index, Gemini 3.8 Flash scored 41 and Gemini 3.1 Pro Preview scored 30 (Artificial Analysis).
  • Main rivals are GPT-6 Astra, GPT-6.1 Sol and Claude Opus/Sonnet 5.5.

Limitations, safety and contested findings

  • The video puts Argon fourth at 57%, behind Sonnet 5.5 (64), Opus 5.5 (60) and Astra (59) (2:24).
  • Internal skepticism: Bloomberg's report of staff doubts about coding is contested by Google (TNW).
  • Staged rollout: Google ties it to feedback and stronger safeguards against cyber misuse (4:25).
Read the full section
  • Terminal-Bench numbers conflict. The video puts Argon fourth at 57%, behind Sonnet 5.5 (64), Opus 5.5 (60) and Astra (59) (2:24). VentureBeat puts Opus 5.5 at 66.4% and Argon last among frontier models (VentureBeat). The two sources also count benchmark wins differently.
  • Internal skepticism: Bloomberg's report of staff doubts about coding is contested by Google (TNW). The video adds that, according to Google, staff tested early versions for weeks (3:02).
  • Staged rollout: Google ties it to feedback and stronger safeguards against cyber misuse (4:25).

Business and practitioner implications

  • Fact-heavy workflows: a 15% hallucination rate favors Argon where a confident wrong answer is costly. Astra's higher accuracy still matters for recall-heavy tasks.
  • Shell and DevOps agents: run your own head-to-head tests. Every source puts Argon behind rivals on terminal work.
  • Cost: model total spend per task, including Argon's ~2× token use and the end of introductory pricing.
Read the full section
  • Fact-heavy workflows: a 15% hallucination rate favors Argon where a confident wrong answer is costly. Astra's higher accuracy still matters for recall-heavy tasks.
  • Shell and DevOps agents: run your own head-to-head tests. Every source puts Argon behind rivals on terminal work.
  • Cost: model total spend per task, including Argon's ~2× token use and the end of introductory pricing.
  • Planning: with no availability date, build evaluation harnesses now and hold off on migrations.
FOLLOW THE EVIDENCE

The source trail.

Sources (9)
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief