OpenAI and Cerebras launch Ultrafast mode: GPT-5.6 Sol at 750 tokens/sec

OpenAI unveiled Ultrafast, an API service tier that runs GPT-5.6 Sol at up to 14x its normal speed, hitting as many as 750 output tokens per second. The acceleration comes via a partnership with Cerebras, whose wafer-scale hardware is purpose-built for high-throughput inference, and OpenAI is billing it as the fastest frontier model available in any API.
The practical target is latency-sensitive and agentic workloads: at 750 tokens/sec, long tool-use loops, code generation and real-time assistants that previously stalled on token throughput become viable. It pairs with OpenAI's concurrent 'builder's guide to GPT-5.6' push encouraging startups to build faster, cheaper agents through smarter model selection and new Responses API capabilities.
The move directly answers a growing speed race — Google's cheaper Flash models and DeepSeek's throughput-tuned harnesses — by competing on tokens-per-second rather than raw intelligence. Community reaction was intense: Cerebras's technical post hit 439 points and 187 comments on Hacker News, with heavy debate over whether the 750 tokens/sec figure holds under real concurrency and what it costs. The open questions are pricing (OpenAI hasn't detailed Ultrafast rates) and how consistently the Cerebras path sustains peak throughput at scale, but the headline capability reframes what 'fast' means for a frontier API.