Google has introduced Gemini 4 Argon, its latest advanced model designed to narrow the performance gap with leading artificial intelligence systems developed by OpenAI and Anthropic. While the model does not claim an outright industry lead, it is notable for its competitive benchmark results and relatively low initial pricing structure.
Model Architecture and Key Features
Argon, Google’s newest frontier model, marks the company’s first such release in over seven months, following the debut of Gemini 3.1 Pro. The model is capable of accepting various inputs, including text, images, video, and audio, but is restricted to generating text outputs. A major technical upgrade is its support for up to one million output tokens, a feature Google calls an industry first. This capability allows the model to process complex reasoning problems in a single sequence without encountering a timeout.
To manage these lengthy outputs, Google is incorporating a new “Long Decode Continuation” feature into the Gemini API. This mechanism pauses extended responses and allows them to be resumed via subsequent requests, ensuring that deep reasoning does not fail due to time limits.
Deployment Schedule and Pricing
Google plans a gradual, phased rollout for Argon. Initially, access will be restricted to “trusted cyber defenders” participating in the Fairwind program, alongside Google’s own internal teams. These early users will receive the model without cyber guardrails. Following this, the company will engage with the US government’s voluntary program, granting agencies pre-public access to new models. Only after gathering feedback from these early testers will Argon be made available to general developers, businesses, and consumers, starting with paying API customers and Google AI Ultra subscribers.
The introductory pricing structure is set at $2 per million input tokens and $10 per million output tokens. The regular pricing is scheduled to increase to $4 for input tokens and $20 for output tokens. Furthermore, cached input tokens benefit from a 95 percent discount, making them highly cost-effective.
Independent and Internal Performance Benchmarks
Independent analysis conducted by Artificial Analysis places Gemini 4 Argon at a “High” reasoning level, where it achieved a score of 53 points on the Artificial Analysis Intelligence Index. This score matches OpenAI’s GPT-6 Astra (max) and Claude Fable 5.1, and surpasses GPT-6.1 Sol (max) by a single point. However, Anthropic’s models remain leaders, with Claude Opus 5.5 scoring 58 points and Claude Sonnet 5.5 scoring 56 points. This represents a 23-point increase compared to Google’s previous frontier model, Gemini 3.1 Pro Preview.
From a cost perspective, at the current promotional rate, performing one Intelligence Index task costs $1.99, which is 60 percent of GPT-6 Astra’s cost of $3.26. After the discount expires, the cost rises to $3.98, placing it about 20 percent higher than GPT-6 Astra. This price advantage is attributed to lower token rates rather than superior efficiency, as Argon uses an average of 62,000 output tokens per task, while GPT-6 Astra requires only 27,000.
Specialized Task Performance
Argon also showed significant improvements in agentic tasks, which had previously been a weakness for Gemini models. On the AutomationBench-AA (Artificial Analysis variant), it achieved first place with a score of 77.5 percent, improving six points over Claude Sonnet 5.5 (max). On the Terminal Bench 4, its score was 57 percent—a 53-point jump over Gemini 3.1 Pro Preview. While this still trails Claude Sonnet 5.5 (64 percent), Claude Opus 5.5 (60 percent), and GPT-6 Astra (59 percent), it marks a substantial advancement.
In testing factual accuracy and handling knowledge gaps (AA-Omniscience), Argon registered a low hallucination rate of 15 percent. Its accuracy, however, reached 50 percent, which is five points below Gemini 3.1 Pro Preview and 13 points below GPT-6 Astra (max) at 63 percent. Overall, Argon scored 42 points, placing it near GPT-6 Astra (43) and GPT-6.1 Sol (42).
Industry and Human Preference Rankings
In Google’s own internal benchmarks, Argon led the Vals Index, scoring 68.9 percent and becoming the first Gemini model to top the index. It placed in the top five across 20 of 22 tested benchmarks, demonstrating particular strength in finance, law, coding, and security. It should be noted, however, that the mid-tier Sonnet 5.5 also outranked Anthropic’s top model Opus 5.5 in this specific test.
When assessed by human raters on Arena.ai, Argon proved to be a strong contender for creative tasks. In the Text Arena, Gemini 4 Argon (High) secured first place with 1,525 points, surpassing Claude Opus 4.6 (High) (second place) by 20 points. The model leads in coding, hard prompts, instruction following, longer queries, and creative writing. Furthermore, it ranked first across all professional fields and for queries submitted in English, Chinese, Russian, and various non-English languages. On a price-to-performance basis according to Arena, Argon is currently the most cost-efficient model at a blended rate of $8 per million tokens.