Google's New Gemini 4 Argon Beats Everyone: Except Where It Doesn't
Google dropped Gemini 4 Argon, and the marketing machine is working overtime. It supposedly crushes GPT-6 Astra and Claude Opus 5.5, but if you look past the shiny charts, it starts to look like another classic case of benchmark-chasing.
Google released Gemini 4 Argon, a model tailored for coding and cybersecurity that allegedly makes OpenAI's GPT-6 Astra and Anthropic's Claude Opus 5.5 look like calculators from the nineties. According to the company's internal data, it takes the top spot in 13 out of 19 categories and ties for first in another.
The headline metric is 77.9% on the DeepSWE v1.1 benchmark, significantly ahead of the competition. Argon also claims to dominate in business process automation and long-context video analysis, which is great if those tests actually reflect how real code gets written on a Friday night.
However, critics are already pointing to the fact that Argon conveniently underperforms in newer tests like Terminal-Bench 4.0 and FrontierSWE v2. It is almost as if the model was trained specifically to pass the tests that Google decided were important this week.
For now, the model remains locked behind a walled garden. Pricing is set at $2 per million input tokens and $10 for output, with a steep increase planned after the initial honeymoon phase. Access is currently limited to a handful of users, leaving everyone else to wonder if the Google AI Ultra badge is worth the impending bill.
Building a model that is a genius in a lab but a middle-manager in real-world scenarios is the modern tech equivalent of a filter that makes your lunch look tastier than it actually is. The race to the top of a spreadsheet is essentially a competition to see who can optimize their math the hardest before the public realizes they are just paying for fancy autocomplete.
Source: Google Blog
Comments
Help shape the next version: Add context or suggest a correction. AI review can add points toward a rewrite. Reviews and updates may take time; a full meter does not guarantee a new version.