OpenAI & Cerebras Clock GPT-5.6 Sol at 750 Tokens/Sec
Pure hardware wizardry just solved the waiting game for deep AI reasoning. Massive silicon wafers are now pumping out complex answers faster than human eyes can physically read, leaving standard graphics cards far behind.
OpenAI and chipmaker Cerebras launched the Ultrafast mode for GPT-5.6 Sol, achieving generation speeds of up to 750 output tokens per second in the API. This delivers a 14-fold speed jump over standard processing without relying on a stripped-down or quantized model.
The system runs entirely on the Cerebras Wafer-Scale Engine, utilizing dinner-plate-sized chips equipped with 44 gigabytes of on-chip SRAM. Model weights stay permanently on the silicon while data moves across wafers, neatly dodging the memory bandwidth bottleneck that usually turns traditional GPU clusters into traffic jams.
During benchmark runs on the 2,500 difficult questions of Humanity’s Last Exam at extreme reasoning levels, GPT-5.6 Sol Ultrafast crossed the finish line in 11 hours and 11 minutes. By comparison, Anthropic's Claude Fable 5 required over 78 hours to finish the exact same benchmark with comparable accuracy.
Standard processing currently costs $5 for inputs and $30 for outputs per million tokens, while the 2.5x Fast mode doubles those rates. The exact price tag for the 14x Ultrafast tier remains undisclosed as access rolls out through a limited preview waitlist.
Silicon scaling is rapidly turning AI latency into an afterthought, leaving developers to ponder how quickly their cloud budgets can evaporate when code runs at the speed of light.
Source: OpenAI
Comments
This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.