rkj dev

AWS Releases Strands Decider 2B Open-Source Decision Model

Fine-tuned from a Qwen3.5-2B base with a custom pointer head, the lightweight model makes sub-second agent decisions and releases under Apache 2.0.

Glowing server hardware in a modern data center symbolizing high-speed computation
Illustration: Low-latency AI decision architectures processing automated workflow tasks.AI-generated illustration

Key takeaways

  • AWS released Strands Decider 2B under an Apache 2.0 license with full weights, training datasets, and reproduction scripts.
  • The 1.9B parameter model swaps standard language-model generation for a 1M parameter pointer head, returning scores and probabilities in a single pass.
  • Benchmarked on JevBench, the reference v19 checkpoint achieved 72.3% accuracy and a 0.342 Brier score, with a 115 ms median latency on an Nvidia RTX 3090.
  • Designed for agentic workflows, the model handles tool verification, query triage, guardrails, and model routing without text-generation overhead.

Amazon Web Services has introduced Strands Decider 2B, a dedicated open-source decision model designed to make AI agent workflows faster and more deterministic. Incubated inside Strands Labs, AWS's experimental agent-development initiative, the model bypasses the computational overhead of text generation to provide direct choices and calibrated confidence scores.

Released under the permissive Apache 2.0 license, the model weights are available on Hugging Face alongside the full codebase, training datasets, and reproduction scripts on the strands-decider GitHub repository. Unlike conversational large language models that write explanations before arriving at a choice, Strands Decider 2B processes state and question schemas in a single forward pass, providing a lightweight classification engine for autonomous software agents.

Abstract digital branches demonstrating categorical choice routing and deterministic evaluation
Illustration: Decision models evaluate structured multi-choice options in a single computational pass.AI-generated illustration

The Shift to System One Decision Models

The release comes amid rapid industry interest in so-called "System One" decision models, an architecture popularized when TypeSafe AI launched its proprietary Jev model in September 2026, as reported by VentureBeat. While traditional large language models generate arbitrary natural language token by token, decision models are restricted to selecting from structured options or rating inputs on defined numerical scales.

According to an announcement on the Strands Agents blog, trading away free-form generation eliminates token-by-token decoding latency and prevents hallucinations of unlisted choices. In exchange, the model delivers fast, bounded judgments alongside probability distributions for tasks that do not require prose synthesis.

Strands Decider 2B answers three distinct schema types through a unified scoring mechanism:

  • Choice questions: Selecting one outcome from a list of supplied options, such as routing a customer inquiry to billing, technical support, or sales.
  • Noul questions: Evaluating binary yes-or-no propositions, such as determining if a user's prompt conveys urgency.
  • Score questions: Placing an input along an ordered rubric or sentiment scale from 0 to 1.

Hobson Architecture and Technical Specifications

Strands Decider 2B contains 1.9 billion parameters and is built using Alibaba's pretrained Qwen3.5-2B base torso. AWS engineers discarded the standard language-modelling head that predicts the next text token and replaced it with a custom "pointer head" containing just over 1 million parameters, a setup described in the project's architecture documentation as "Hobson."

A software engineer analyzing neural network configurations at an office workstation
Illustration: Developers integrating lightweight pointer-head architectures into local agent pipelines.AI-generated illustration

According to the project's GitHub repository documentation, the Qwen torso is adapted with a rank-16 LoRA (low-rank adaptation) update while the pointer head executes in fp32. Instead of maintaining fixed per-option output weights, the pointer head calculates an attention score between the hidden state at the <answer> token position and the hidden state at each supplied option's final token.

Because options are scored dynamically through this masked softmax mechanism, the model places no architectural ceiling on the number of options provided in a prompt. The current release is designated v19, representing an architectural shift away from an earlier "slot head" design that AWS researchers found performed significantly worse.

Benchmark Performance and Latency

AWS evaluated the reference v19 checkpoint on the public 231-task suite of JevBench v1. On this benchmark, Strands Decider 2B achieved an overall accuracy of 72.3% (167 of 231 tasks) and a Brier calibration score of 0.342, with an expected calibration error of 0.052.

In split evaluations published in the repository documentation, the model answered 100% of tasks in JevBench's easy tier correctly, alongside 87.5% on standard tasks and 50.5% on hard tasks. On short classification tasks, predictions output with a confidence score of 0.90 or higher were correct approximately 95% of the time.

Hardware benchmarks demonstrate low operational latency across both server and consumer silicon:

  • Nvidia RTX 3090: Measured across 230 requests, the model posted a median decision latency of 106 ms to 115 ms and a 95th-percentile (p95) latency of 296 ms to 299 ms.
  • Apple Silicon (M3 Pro): Under warm serving conditions for small tasks under 300 tokens, the model registered a median latency of 153 ms.
  • Training Speed: The full training recipe runs in approximately 11 hours on a single 24 GiB RTX 3090, or 1 hour and 10 minutes on an 8-GPU Nvidia H100 node.

Competitive tracking reported by VentureBeat indicates that while Strands Decider ranks near the top of its size class, it trails Mapika's decider-2b v11, which scored 76% accuracy and a 0.32 Brier score on the same public JevBench set. TypeSafe's proprietary Jev was not plotted directly on the same open weight trajectory.

Deployment in Agentic Interventions

AWS designed Strands Decider 2B to integrate directly with agent development frameworks, including the Strands Agents SDK. Rather than querying costly frontier models to approve intermediate steps, developers can place the 2B model inside synchronous intervention hooks.

In a demonstration published by AWS, Strands Decider acts as an automated review gate before an agent executes a weather tool call. If a user asks for local weather without specifying a city, eager agents often guess a location. Decider inspects the conversation history and answers two binary questions: whether the tool arguments are grounded in the user's actual text, and whether calling the tool is premature. Based on the model's confidence scores, the application denies the immediate tool execution and instructs the agent to ask the user for clarification.

Beyond tool gating, AWS identified model routing, policy enforcement, argument checking, output evaluations, and triage queues as primary deployment scenarios. Pairing lightweight decision models with generative engines enables hybrid agent architectures where small models handle rote execution checks and frontier models perform unstructured reasoning.

Availability and Commercial Context

Strands Decider 2B is completely free to download and self-host, but AWS has not launched a managed cloud API endpoint or published general per-request hosted pricing. In contrast, TypeSafe AI offers Jev as a managed API priced at $0.042 per million input tokens with zero output token fees.

The launch also coincides with releases from other infrastructure providers, including Cloudflare's open-source 27-billion-parameter Clef model on Hugging Face. While self-hosting requires teams to manage local GPUs or cloud compute infrastructure, the complete release of training data and scripts allows organizations to inspect, fine-tune, and adapt the decision architecture to proprietary data entirely on premise.

Frequently asked questions

What is Strands Decider 2B?

Strands Decider 2B is an open-source, 1.9-billion-parameter decision model released by AWS Strands Labs that selects options and scores inputs without generating free-form text.

How does Strands Decider 2B differ from traditional LLMs?

Instead of predicting sequential text tokens, it replaces the language model head with a 1-million-parameter pointer head that scores predefined choices in a single forward pass, providing calibrated confidence probabilities.

What hardware is required to run the model?

Strands Decider 2B can be served locally on consumer GPUs such as an Nvidia RTX 3090 with roughly 115 ms median latency, as well as on Apple Silicon Macs (M3 Pro) and standard CPUs.

Under what license is Strands Decider 2B available?

The model weights, code, and full training recipes are published under the open-source Apache 2.0 license.

Sources

  1. Cloudflare/clefHugging Face · Sep 30, 2026 · Official
  2. strands-labs/strands-deciderGitHub · Sep 29, 2026 · Official
  3. Amazon unveils a free, fast, open source Jev killer: Strands Decider 2B makes decisions in fractions of a secondventurebeat.com · Oct 1, 2026
  4. AWS debuts Strands Decider 2B, a first lightweight decision model for accelerate agentic workflowsSiliconANGLE · Oct 1, 2026
  5. Introducing Strands Decider 2B: a small, open source, decision modelStrands Agents · Oct 1, 2026
  6. AWS launches a local answer to TypeSafe’s Jev decision modelThe New Stack · Oct 1, 2026
  7. Amazon Ships Strands Decider 2B, an Open-Source Jev RivalAI Weekly · Oct 1, 2026

How this story was made: the newsroom picked it up from siliconangle.com, Reddit and Techmeme, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (38 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.

#AWS #Open Source #AI Agents #Qwen #Machine Learning

Published October 2, 2026 at 01:08 UTC