Back
OpenAIAugust 13, 20261 sources

OpenAI launches GPT-5.6 Sol 'Ultrafast' mode at 750 tokens/second on Cerebras

AI Analysis

OpenAI launched 'Ultrafast,' a new operating mode for GPT-5.6 Sol that delivers about 750 output tokens per second — roughly 14 times faster than the standard mode — running on Cerebras wafer-scale chips rather than GPUs. The mode is in limited preview and aimed squarely at latency-sensitive enterprise use cases: incident response, customer service, financial analysis, and e-commerce.

The significance is architectural. By routing a frontier model onto Cerebras silicon, OpenAI signals it will use non-NVIDIA accelerators where they win on throughput, and that inference speed — not just intelligence — is now a product axis worth a distinct SKU. For agentic workflows that chain many model calls, a 14x speedup can be the difference between a usable real-time agent and an unusable one.

The launch came amid continued executive turnover at OpenAI: the company replaced CRO Denise Dresser with Wiz President Dali Rajic, and separately shipped consumer features including 'Computer History' (ChatGPT remembering Mac activity, rolling out to Pro/Business/Enterprise) and reservation search teased by Greg Brockman. Those are distinct announcements; Ultrafast is the enterprise-facing headline.

Competitively, Ultrafast puts OpenAI in direct contrast with Groq and Cerebras-hosted rivals that have marketed extreme token throughput, and against Grok 4.6 and Gemini 3.7 Flash's speed-and-cost positioning this same week. The caveat is availability — it's a limited preview with unstated pricing and capacity, and Cerebras supply is finite. Watch whether Ultrafast graduates to general availability and how its per-token economics compare to standard GPU-served Sol.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog