OpenAI's Ultrafast mode runs GPT-5.6 Sol at 750 tokens/sec via Cerebras

OpenAI has introduced Ultrafast, a mode that runs GPT-5.6 Sol at up to 750 output tokens per second — about 14x the speed of standard GPT-5.6 Sol processing — powered by a partnership with specialized AI-chip maker Cerebras. The speed leap targets latency-sensitive agentic workflows where a model must run many sequential steps: at 14x throughput, long agent chains that felt sluggish become interactive, and real-time coding assistance gets dramatically snappier.
The mechanism is Cerebras' wafer-scale hardware, which keeps model weights on a single massive chip and avoids the memory-bandwidth bottlenecks that cap token throughput on conventional GPU inference. Cerebras has built its business on this speed advantage, and the OpenAI partnership is a marquee validation. Cerebras' post detailing the collaboration drew 698 points and 272 comments on Hacker News, one of the day's top discussions on inference speed.
Competitively, Ultrafast is OpenAI's answer to a growing emphasis on inference speed as a differentiator — Google's Gemini Flash line and various providers compete on cost-per-token, but raw latency is an underexploited axis. For enterprise users running agents at scale, 14x speed can translate directly into throughput and lower effective cost per completed task.
The caveats, raised on Reddit: Ultrafast is a limited preview and offers no open-weight alternative, so it deepens dependence on OpenAI's proprietary stack plus Cerebras hardware. There's also the question of whether 750 tokens/sec is sustained or peak, and how pricing compares to standard Sol. The reliance on Cerebras — a much smaller supplier than NVIDIA — raises capacity questions if demand surges. Watch pricing details and whether the preview widens to general availability.