Anthropic Drops Claude Sonnet 5.5: Crushing Opus at Half the Price
The mid-tier officially declared war on the ultra-expensive flagships, but keeping your cloud billing account intact will require reading the very fine print.
Anthropic has rolled out Claude Sonnet 5.5, locking in baseline token pricing at two dollars per million input tokens and ten dollars per million output tokens—neatly cutting the cost of the flagship Opus 5.5 straight in half.
On the command-line programming benchmark Terminal-Bench 4.0, Sonnet 5.5 racked up a 70.6% accuracy rate, easily overtaking Opus 5.5 at 66.4%. Across desktop control benchmarks like OSWorld 2.1, it virtually mirrored the flagship with an 80.1% score. The independent Intelligence Index by Artificial Analysis ranked the new model second globally at 56 points, leaving OpenAI's GPT-6 Astra sitting in third place.
The financial hangover strikes only when engineers dial the model up to full throttle. Independent tests revealed Sonnet 5.5 burns through roughly 193,000 output tokens per task at maximum effort, jacking the single-run price tag up to $7.60—nearly 50% steeper than its predecessor. While moderate settings deliver genuine savings, the maximum setting treats reasoning tokens like complimentary appetizers.
The underlying knowledge base exposes a clear operational profile: the model excels at doing work rather than memorizing trivia. It notched just 54% on the AA-Omniscience factual test compared to 66% for Opus 5.5, though it gained frontier-grade anti-distillation barriers that block rivals from harvesting its internal reasoning traces.
Flagship AI models are rapidly turning into museum pieces reserved for exotic edge cases, while cheaper alternatives handle the actual day-to-day execution. The premier engineering skill of this cycle will not be crafting clever prompts, but aggressively clamping down the reasoning slider before the accounting team sounds the alarm.
Source: Anthropic
Comments
Help shape the next version: Add context or suggest a correction. AI review can add points toward a rewrite. Reviews and updates may take time; a full meter does not guarantee a new version.