GPT-6 Astra Ultrafast runs up to 8x faster on NVIDIA Blackwell, reaching the API and Bedrock

Even as OpenAI shelved the next Astra, the current one got much faster. NVIDIA's blog says GPT-6 Astra Ultrafast now runs in the OpenAI API and for eligible ChatGPT Work and Codex users, generating tokens up to 8x faster than Astra Standard. AWS separately listed an UltraFast premium tier for GPT-6 Astra on Bedrock.
The speedup comes from inference optimizations built for Blackwell, NVIDIA's current GPU architecture. These likely include low-precision kernels, speculative decoding and memory-bandwidth tuning that only pay off on that hardware. The result is a speed tier rather than a new model: the same weights, served faster at a premium. For Codex users the gain matters in practice, because agentic coding loops chain many generations and waiting time compounds across them.
This is the latest round in a speed-tier race. Anthropic and Google offer fast modes, and specialist inference providers such as Cerebras and Groq built their pitch on raw tokens per second. OpenAI tying Ultrafast to Blackwell also strengthens its NVIDIA relationship in a week when OpenAI was notably absent from NVIDIA's agent-safety coalition. The relationship is commercial even where the companies differ on governance.
Caveats: '8x' is a peak figure relative to Standard mode, not a guaranteed throughput, and pricing for the premium tier is not given in the sources. Developers on HN already resent tightened ChatGPT Pro quotas, so a pay-more-for-speed tier may sharpen the complaint that OpenAI is monetizing access rather than capability. Watch for real tokens-per-second measurements and Ultrafast versions of GPT-6.1 Sol.