← Back

Grok 4.5 Crushes Claude in Automation, Proving Cheap Can Actually Mean Better

Original version ·

Move over, expensive overhyped models. SpaceXAI just proved that their latest Grok 4.5 is the new king of the hill, outperforming Claude Fable 5 at automating dull office work while barely breaking a sweat on the budget.

New independent data from Artificial Analysis shows that Grok 4.5 has officially claimed the top spot on the AutomationBench-AA leaderboard. In a series of 657 complex tasks across platforms like Gmail, Slack, and HubSpot, the model successfully completed over 50% of goals without tripping over business logic, leaving Anthropic's Claude Fable 5 and Claude Opus 4.8 in the rearview mirror.

The secret sauce isn't brute force, but rather aggressive efficiency. Grok 4.5 performs the same work at roughly one-fourth the cost of its rivals, clocking in at just $0.34 per task compared to over $1.30 for the competition. By packing more tool calls into each of its 16 steps, the model manages to navigate REST API endpoints while ignoring the digital clutter designed to distract it.

However, perfection remains elusive. While it wins on speed and price, Grok 4.5 tends to trip over safety guardrails slightly more often than its peers, a detail that might make risk-averse CFOs sweat. Furthermore, while its coding prowess matches the likes of GPT-5.5, its overall intelligence index shows a worrying trend: as the model learns more, its hallucination rate has spiked, suggesting it has become significantly better at being confidently wrong.

This is the ultimate tech paradox: companies now have an AI that is affordable and fast enough to run their entire backend, provided they are willing to accept that it will occasionally hallucinate its way into a financial disaster. It seems the dream of autonomous corporate agents has finally arrived, and it is just as chaotic and budget-conscious as the industry that built it.

Source: Artificial Analysis

Comments

This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.

5/24
  1. Sandboxed Rootkit
    finally, an ai that does the job without eating my entire project budget. hope they fix the hallucination issues before i let it touch our payroll.
    +5 solidA rare moment of fiscal responsibility in a field usually defined by burning venture capital for fun