Skip to content
The AI-First Web

Google's Gemini 4 Argon ties GPT-6 Astra on score but guesses far less often

Artificial Analysis finds a 15% vs 51% hallucination rate. The price edge, though, rests on a launch discount.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Google's Gemini 4 Argon ties GPT-6 Astra on score but guesses far less often

Photo: Google DeepMind / Pexels

Key finding

Argon hallucination rate (AA-Omniscience): 15% (Source: Artificial Analysis (September 30, 2026))

A tie on the score hides a difference in how models fail

Two AI models can earn the same grade and still behave very differently at work. One admits when it does not know. The other fills the gap with a confident guess. Most leaders would hire the first. A single headline score does not show which one they are buying.

That is the lesson in the latest results from Artificial Analysis, an AI benchmarking firm. Google's Gemini 4 Argon, run at its highest reasoning setting, scores 53 on the firm's Intelligence Index. That matches OpenAI's GPT-6 Astra (max) at 53 and edges GPT-6.1 Sol (max) at 52.

What the researchers found

On the same index, the models look equal. On a second test, they do not. Artificial Analysis measured how often each model guesses wrong instead of admitting it does not know. It calls this the hallucination rate.

15%
Argon hallucination rate (AA-Omniscience)
Source: Artificial Analysis (September 30, 2026)
51%
GPT-6 Astra (max) hallucination rate
Source: Artificial Analysis (September 30, 2026)

Artificial Analysis says Argon's rate is the lowest of any model scoring 45 or more on its index. GPT-6.1 Sol (max) sits at 54%. In the report's words, Argon is much more likely to acknowledge it does not know an answer than to guess wrong.

The trade-off behind the number

Admitting ignorance has a cost. Argon's accuracy on the same test is 50%. That is 13 points below Astra's 63% and 5 points below Google's previous Pro model.

So Argon answers fewer questions correctly. But a much lower share of its non-correct responses are confident guesses. Netted together, the overall AA-Omniscience scores land close: 42 for Argon, 43 for Astra and 42 for Sol.

This is where an average misleads. Think of two new hires with the same performance review score. One says "let me check" when unsure. The other improvises. Which one you want depends on the work. A drafting assistant can tolerate a bad guess that a person will catch. An automated workflow that sends a customer reply or updates a record cannot. That is our interpretation, not a finding from the report.

The report also points to stronger agentic results for Argon, meaning tasks where a model takes steps on its own. It ranks first on AutomationBench-AA at 78%. On Terminal Bench 4 it scores 57%, behind only two Claude models and GPT-6 Astra (59%). The top score there is 64%. These are Artificial Analysis's own tests, not results on your workloads.

The price is a promotion

The cost story needs the same care. Google is selling Argon at half price during its rollout. At that rate, one Intelligence Index task costs $1.99, against $3.26 for Astra. That is about 60% of Astra's bill. At standard pricing the figure becomes $3.98, roughly 1.2 times Astra.

$1.99
Argon cost per task, at launch discount
Source: Artificial Analysis (September 30, 2026)

Google has not confirmed when the discount ends. The report calls it an initial promotion that runs for at least one month. Argon is also rolling out to selected users only and is not publicly available.

The saving comes from lower token prices, not from efficiency. Tokens are the small chunks of text a model reads and writes, and they set the bill. Argon averaged 62,000 output tokens per task against 27,000 for Astra. Compared with GPT-6.1 Sol (max), Argon at discount costs 2.7 times as much per task.

One new feature matters for cost planning. Long Decode Continuation lets a long response pause and resume across follow-up calls. Artificial Analysis says this allows reasoning to run up to 1 million output tokens without request timeouts. Long runs mean more tokens, and more tokens mean a bigger bill.

What to ask your team

Do not choose a model on one score. Put these questions to the people running your AI projects:

1. Which costs us more on this task: a wrong answer delivered with confidence, or a correct answer withheld? Test models against that answer, not against a leaderboard.

2. Do we measure how often our model declines to answer? If not, we cannot tell a cautious model from a weak one.

3. Is any price in our business case a promotional rate? Rebuild the figure at the standard rate of $4 per million input tokens and $20 per million output tokens.

4. How many tokens does a real task use? Per-token price and per-task cost can point in different directions, as the Argon and Astra gap shows.

A tie on intelligence is only the start of the comparison. The question for buyers is how the model behaves when it does not know.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Artificial Analysis.

Share this insight