Claude Opus 5 lands: 30.2% on ARC-AGI-3, half the cost of Fable 5

Anthropic positioned Claude Opus 5 as 'thoughtful and proactive,' delivering greatly improved performance at the same price point as its predecessor Opus 4.8 while approaching the frontier intelligence of Claude Fable 5 at half the price. The headline number is ARC-AGI-3: Opus 5 scored 30.2%, versus 7.8% for OpenAI's GPT-5.6 Sol (Max) — a benchmark explicitly designed to measure general reasoning rather than memorized skills. Pricing lands at $5 per million input tokens and $25 per million output tokens, and Opus 5 becomes the default on Claude Max and the strongest option on Claude Pro.
AWS's Swami Sivasubramanian, announcing availability on Bedrock, emphasized long-running agent behavior: Opus 5 'powers long-running agents that work for hours and even overnight,' navigating codebases 'like an experienced engineer,' pushing back on flawed instructions rather than executing blindly, and spawning sub-agents that need less supervision.
The developer verdict is more divided. On r/Anthropic, a thread arguing Opus 5 'feels so much more painful than 4.8' drew 118 comments, and a second ('Opus 5 is really testing my patience') hit 183 upvotes — suggesting the model may expect different prompting conventions. LlamaIndex's Jerry Liu benchmarked it on document parsing (ParseBench) and found it 'roughly on par with Opus 4.8,' slightly worse on dense tables, and — at 8 cents per page — not worth using for large-scale OCR versus Gemini 3.6 Flash or LlamaParse. His verdict: great for coding and knowledge work, poor value for document parsing at scale.
Separately, Anthropic faced a privacy embarrassment as Claude shared chats were reportedly indexed by Google Search. The competitive frame is stark: Opus 5's benchmark lead over GPT-5.6 Sol arrives just as cheap Chinese open models crowd the market from below.