DeepSeek V4 hits general availability with 1M-token context and peak/off-peak surge pricing

DeepSeek V4 reached general availability, standardizing a 1M-token context window across all tiers and adding improvements in agentic execution, math reasoning, and code generation. The launch carries two operationally significant changes for API users. First, DeepSeek becomes the first frontier-tier model API to formalize surge pricing: rates roughly double during Beijing business hours (peak) versus off-peak, an explicitly cost-aware but friction-inducing structure for teams outside China's time zone. Second, a hard migration deadline — legacy aliases deepseek-chat and deepseek-reasoner retire July 24 at 15:59 UTC with zero grace period — forces users onto deepseek-v4-flash or deepseek-v4-pro.
The zero-grace-period cutoff is the developer pain point, flagged across communities as an abrupt break that will strand un-migrated pipelines. The peak/off-peak model, meanwhile, is being read as a template others may copy to manage GPU demand.
Competitively, DeepSeek slots into the same China open-weight momentum as Qwen 3.8 and Kimi K3, and r/DeepSeek reaction is strongly positive: 'Deepseek V4 Pro is AMAZING' (383 upvotes) and threads pairing DeepSeek with Claude Code as 'the most cost efficient agentic workflow for dev.' Pre-launch chatter pegged V4 as close to GPT-5.6 Sol. The affordability and 1M context are the draw; the surge pricing and migration deadline are the caveats. This also intersects with Microsoft's reported interest in Chinese models for Copilot. Watch adoption after the July 24 cutover and whether U.S. teams tolerate time-zone-dependent pricing.