Back
OpenAIAugust 03, 20261 sources

OpenAI GPT-5.6 models on AWS Bedrock now support 1M-token context, get up to 80% price cut

AI Analysis

Amazon Web Services expanded the context window for OpenAI's GPT-5.6 Sol, Terra, and Luna models on its Bedrock platform to 1 million tokens, letting enterprises process massive data volumes — such as entire codebases or document corpora — in a single inference call. Prompt caching offers a 90% discount on repeated context, and AWS aligned the on-Bedrock pricing with OpenAI's first-party rates.

In its weekly roundup, AWS announced price reductions of up to 80% for the GPT-5.6 family through Amazon Bedrock, a significant move that lowers the barrier for high-volume enterprise deployments. The cut mirrors OpenAI's own aggressive repricing — the company recently slashed GPT-5.6 Luna API prices 80% to $0.20 input / $1.20 output per million tokens on its first-party platform.

The expansion is part of AWS's Q2 momentum: the cloud unit posted 36.7% YoY growth, its fastest in 18 quarters, with a $496 billion backlog, and raised FY26 capex guidance to roughly $220 billion. Analysts reframed that capex hike as constraint-driven — Amazon paying elevated prices to protect its GPU and HBM allocation queue — rather than pure confidence. The 1M-context-plus-price-cut combination directly targets enterprises weighing whether to run frontier OpenAI models inside their governed AWS environment versus calling OpenAI directly. Watch whether latency and throughput at 1M tokens hold up under real production loads.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog