Skip to main content
The Quantum Dispatch
Back to Home
Cover illustration for Strands Decider 2B: What Amazon's Free Decision Model Does

Strands Decider 2B: What Amazon's Free Decision Model Does

Strands Decider 2B is a free 2B-parameter decision model that picks options in about 115ms on a single GPU. Here is how it works and where it fits.

Dr. Nova Chen
Dr. Nova Chen★Oct 3, 2026★4 min read

Strands Labs, the Amazon Web Services team behind the open-source Strands Agents SDK, released Strands Decider 2B on October 1, 2026. It is a small open model built for one job: choosing. Instead of writing paragraphs, Strands Decider 2B looks at a fixed list of options and returns a calibrated confidence score for each one, which makes it a fast, cheap building block for AI agents.

  • Size: 2 billion parameters, built on Qwen 3.5-2B with a rank-16 LoRA adapter.
  • Speed: a median of around 115ms per decision on an Nvidia RTX 3090 and about 153ms on an M3 MacBook, per Strands Labs.
  • Accuracy: 3rd of 33 models in the 2B class on JevBench, and 100% on the benchmark's easy tasks.
  • Availability: free, with code on GitHub, weights on Hugging Face and a one-line pip install.

What Is a Decision Model?

Most language models generate text one token at a time. A decision model skips that. You hand it a question and a short list of possible answers, and it scores every option at once. The output is a probability for each choice, so your software knows not just what the model picked but how sure it was.

That idea went mainstream last month with TypeSafe's Jev, which we covered in our explainer on Jev's probability-based answers. TechCrunch describes Strands Decider 2B as Amazon's own take on the format, arriving as several decision models land in quick succession.

How Strands Decider 2B Works Under the Hood

According to the Strands Labs blog, the team removed the base model's language-modeling head, the part that predicts the next word, and replaced it with a pointer head. The pointer head compares the model's internal state at each option against its state at the answer position, then scores the match. A small rank-16 LoRA adapter tunes the Qwen 3.5-2B base for the task.

The result is a model that is cheap to run. Strands Labs says it works on a local CPU or GPU, and the published latency figures put a single decision well under a fifth of a second on consumer hardware.

Where Would You Use a Decision Model?

Strands Labs lists a long menu of jobs. The most useful for teams building agents today:

  • Model routing: decide whether a request needs a big frontier model or a cheap small one.
  • Tool selection: pick which tool an agent should call next.
  • Guardrails and policy checks: classify a request against a fixed set of rules.
  • Memory and context management: choose what to keep and what to drop.
  • Evaluations: grade outputs against a rubric with calibrated scores.

Each of these is a multiple-choice problem that teams often hand to a large chat model today. Moving them to a 2B-parameter local model can cut both latency and cost, and it keeps routine decisions on your own hardware.

Is Strands Decider 2B Good Enough for Production?

The benchmark numbers are promising for its size. Strands Labs reports a 3rd-place finish out of 33 models in the 2B class on JevBench, and first of 30 once you exclude models just over 2 billion parameters. Those are the vendor's own results, so treat them as a starting point and test on your own decisions before relying on it.

The low barrier to trying it is the real story. One pip install gets you running, and the open weights mean you can fine-tune it on your own options. For anyone building open-weight AI agents, a fast, free chooser is a welcome new part. Follow more launches like this in our AI coverage.

Sources: Strands Agents Blog — October 1, 2026; TechCrunch — October 1, 2026.

More AI Stories