AWS open-sources Strands Decider 2B, a 115ms decision model for agent routing

AWS has joined the push toward small 'decision models.' Strands Decider 2B is a 2-billion-parameter open-source model built on the Qwen3.5-2B base. It does not generate free-form text. Instead, it selects an answer from a list of options and attaches a confidence score, returning results in about 115 ms on a consumer RTX 3090.
Swami Sivasubramanian, AWS VP of Agentic AI, said decision models 'always produce an answer from the selected options' and run with very low latency on a local CPU or GPU. AWS Chief AI & Technology Officer Matt Wood listed the target uses:
- which tool to call
- which model to route a request to
- whether an action passes a guardrail
- simple classification such as spam detection
These are the small decisions an agent makes on every turn. Today, many teams send them to an expensive frontier model.
The category is filling up quickly. Perplexity open-sourced pplx-decider-27b with a Decisions API at 4 cents per million input tokens, and Hugging Face added decision-model support to llama.cpp. Building on Qwen also raises a question this week, since independent audits found censorship embedded in Qwen weights. Enterprises may want to know what a Qwen base carries into a routing model used for guardrail decisions.
AWS also previewed the Well-Architected Agent, which uses AI to optimize cloud environments for cost, security and performance. Watch for Strands Decider integration into AgentCore and for whether independent benchmarks compare it with Perplexity's larger model.