OpenAI Slashes GPT-5.6 Prices Up To 80% And Boosts Sol Speed
The AI pricing war is spiraling out of control as OpenAI frantically edits its fee menu again. When developers start complaining about token bills, tech giants suddenly remember how to offer discounts.
The newest pricing shift hits the GPT-5.6 family, where the budget-friendly Luna variant sees an 80% price slash down to $0.20 per million input tokens. Meanwhile, the mid-tier Terra model gets a 20% discount, dropping its API baseline to $2 per million input tokens while output sits at $12. The corporate bean counters seemingly realized that burning money on inference is only fun when users do not see the invoice.
Subscribers on Codex and ChatGPT Work plans will not see lower monthly subscription fees, but Luna and Terra will simply consume fewer plan credits per request. On the performance front, OpenAI left the price of flagship model Sol untouched, choosing instead to attach a new Fast mode toggle that replaces the legacy Priority Processing tier. This feature delivers up to 2.5x faster response speeds at double the price, reachable via the service_tier: "fast" API parameter without compromising output quality.
To push latency even lower, select preview customers are now running Sol directly on hardware from Cerebras, reaching speeds of up to 750 tokens per second.
Slashing prices while charging double for speed proves that latency is the new tax on modern AI engineering. Wall Street will eventually demand profits, but today cheap cloud computing keeps developer loyalty locked in for another quarter.
Comments
This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.