AI news, with the context that matters.

Technology & Research · ·

Strands Decider 2B is a small AI model for making choices locally

A large LLM plans the work while a small model handles recurring choices. We explain how Strands Decider 2B works, how it can route tasks to local specialists, what confidence means, and the conditions behind its hardware support and performance figures.

Applications built around large LLMs make many small choices between writing tasks: which team should receive a request, which tool should run next, or whether to ask the user before proceeding. Strands Decider 2B is a small AI model designed to handle those choices locally.

The Strands team released v19 on October 1, 2026. Give the model a situation, questions, and options, and it returns numerical results for the options. It does not generate prose, code, summaries, or tool arguments. It fits a division of work in which a large LLM plans and writes the final answer while a smaller model handles recurring classification along the way. Official announcement

This article is based on public documentation and code checked on October 4, 2026, Japan time. Performance figures are the developers' measurements. LATENT has not installed the dependencies, downloaded the weights, run the model, or conducted benchmarks. The routing diagram below is an illustrative architecture that we have not executed.

Key takeaways

  1. Strands Decider 2B is a small model for local use that scores and selects from supplied options. It does not generate prose, code, or tool arguments.
  2. It suits a workflow in which a large LLM plans and writes the final answer while Decider assists with recurring routing decisions. The application controls actual tool execution and requests for human confirmation.
  3. For choice, confidence is derived from the highest probability and the number of options; it guarantees neither accuracy nor safety. Calibration may not transfer to long documents or other uses, so each use case needs evaluation.

Option scoring replaces the text generation head

Decider uses the language-processing torso of Qwen3.5-2B-Base. The developers removed the LM head that predicts the next token and attached a pointer head that scores options.

Generation and selection produce output differently

A typical generative LLM

Input

Context and instructions

↓
Model

Language-processing torso

LM head predicts the next token

↓
Output

Generates successive tokens

token 1→token 2→ …

Prose, code, structured data, and more

Decider choice

Input

state + question + supplied options

billingsalesretail
↓
Model

Qwen3.5-2B-Base language torso

Pointer head scores the options

↓
Output

Selects from the supplied options

Selected option identifier
Per-option probabilities and confidence

LATENT diagram based on the official architecture documentation. The Decider side shows choice. Options are supplied with each request; it generates neither answers outside that set nor text.

The pointer head compares the internal representation at the end of each option with the representation at the answer position to calculate its score. There is no loop generating one token after another.

The pointer head has about a million parameters, and the torso is adapted with rank-16 LoRA. LoRA trains a small set of additional parameters instead of updating all the original weights. The developers describe the base model as having about 1.9 billion parameters.

Each request supplies descriptions of what its options mean. Adding a department therefore does not require rebuilding a classification output layer. Model architecture

The input state holds text or structured data, and questions specifies the questions and options. The same underlying computation is read back differently for each question type.

TypeOutput
choiceThe selected option, probabilities for all options, and confidence. The API accepts 2–255 options.
noulP(true) for a yes/no question. There is no separate confidence field.
scoreAn expected value, probability distribution, and confidence over an ordered scale of 2–10 levels.

For example, on a three-level scale of 0, 1, and 2, score can return an average position such as 1.10. It does not produce an explanation in free text. And if the correct answer is missing from the options, selecting within that set cannot create it. Input and output definitions

A large LLM plans while a small model handles routine routing

For a single question in an ordinary chat, a large LLM can make the judgment and write the answer together. Decider becomes more useful when an application repeats the same kind of decision. Avoiding a large-model call for each routing step can reduce API calls and waiting time. Local inference still has compute and maintenance costs.

Consider a large LLM planning an answer from several internal documents. For each task, the application offers four choices: an extraction specialist, a translation specialist, a return to the large LLM, or a question for the user. It sends the task material, required output, and results so far to Decider, then maps the selected option to an actual operation.

Decider selects; the application decides what runs

Large LLM · may run in the cloud

Interprets the request and plans the work

↓ Pass task state and permitted options

This part runs locally

1 · Decider selects

Chooses extraction, translation, the large LLM, or user confirmation

Returns the option identifier and numbers

→
2 · Application checks

Checks missing input, task-specific thresholds, and execution rules

The model's choice alone does not trigger execution

Branch A · Conditions are met

↓

Application executes the task

Local specialists or tools
extract fields, translate, and more

↓ Return task results

Branch B · Uncertain, incomplete, or needs confirmation

↓

Hold execution and refer back

Large LLM reconsiders
or the user is asked to confirm

↓ Return reconsideration or confirmation results

Large LLM · may run in the cloud

Checks results, revises the plan, or writes the final answer

An unexecuted LATENT example based on the project's suggested uses and division of work. It is neither an official product architecture nor a safety guarantee. Data passed to a cloud LLM leaves the device.

Decider returns an identifier from the supplied options and numerical results. The application or large LLM must define the fields to extract, write translation instructions, and construct tool arguments. The specialist models also need to be supplied separately; installing Decider does not provide that entire system.

Complex changes to the plan and synthesis across results remain with the large LLM. The application also defines the conditions for referring a task back when setting its execution rules.

Whether the entire workflow stays local depends on where each model runs and which data the application passes to it.

Confidence is not a direct statement of correctness

For choice, confidence is calculated from the largest probability and the number of options. In the official CLI example, billing's probability of 0.845 and confidence of 0.768 are related by the formula below. Confidence is not produced by a separate predictor.

Confidence is calculated from the option probabilities

Official CLI team-routing example · billing is selected

billing0.845
retail0.091
sales0.064
01

Each bar's length shows that option's probability.

Derived from the maximum and option count

confidence 0.768

Option count N = 3
Maximum probability p_max = 0.845

(N × p_max − 1) / (N − 1)

(3 × 0.845 − 1) / 2
= 0.7675 ≈ 0.768

This is not a measure of accuracy.

Chart of the official announcement's example output, alongside the choice formula. LATENT did not run this example. Using the displayed probability gives 0.7675, or 0.768 rounded to three decimal places.

Confidence maps a uniform distribution to 0 and all probability on one option to 1. A value of 0.768 cannot simply be read as a 76.8% chance of being correct. The 0.845 is also a model probability conditional on the supplied situation and options, not a guarantee of real-world accuracy or safety.

For score, confidence uses a different formula. It measures the standard deviation of the distribution over the scale and corrects for the training process that spreads some target probability to adjacent levels. noul returns only the probability of true. Similar names do not make these numbers interchangeable. Confidence implementation

The released model fits a temperature for each question type using short classification tasks held out from training. This adjusts overconfident or underconfident probabilities. The developers report that the adjustment transfers poorly to long documents and unfamiliar scales. Documentation examples such as automatic action above 0.9 are not universal thresholds. An application needs to measure how often its own inputs produce decisions that should have been allowed or stopped. Limits of calibration and transfer

Results on 231 public tasks do not measure overall workflow success

The developers use JevBench's public set of 231 tasks. It compares predictions with specified answer labels across 18 families of choice-based tasks. The benchmark comes from a third party, but the results presented here were run by the Strands developers, not independently reproduced by LATENT.

MeasurementCorrect tasks and Brier
Released Hugging Face weights
4,096-token window
167 / 231
Brier 0.348
3,072-token window in the same model card167 / 231
Brier 0.349
Original GitHub v19 evaluation
3,072-token window
167 / 231
Brier 0.342

167 / 231 is approximately 72.3%. The public model card and the original GitHub experiment report different probability metrics even though both are labeled v19. GitHub also compares a retraining run and reports 168 / 231 for the original experiment at a 4,096-token window. That 168 should not replace the published weights' result. The model card further notes that its 3,072-token run used a copy with symlinked weights whose hashes were not checked. Released model card / Original and retraining results

Accuracy in this evaluation is the fraction of tasks for which the highest-probability label is correct. Scale questions also use the top label for accuracy; expected-value error is handled separately. Correctness scoring

Brier score assigns the correct option a target of 1 and every other option a target of 0, then squares the difference from each predicted probability. These errors are summed over the options and averaged across tasks. Lower is better, and binary questions also sum the errors for both outcomes. It is neither a simple error rate nor a measure isolating calibration alone. Brier and other metric definitions

In the developers' difficulty split of the original v19 experiment, easy scored 48 / 48, standard 63 / 72, and hard 56 / 111. The claim of 100% on easy tasks is bounded by those 48 questions. It does not establish that the model can reliably solve long conditional passages or multi-step reasoning. Public-set accuracy is also separate from JevBench's composite score, which includes sealed tasks, speed, and cost. Public-set conditions

The announcement's third-place claim in the 2B class also depends on the public-task results and model-size grouping used at the time. The GitHub record clarifies that v19 itself was not listed on the leaderboard. Retraining also changes results, making a difference of a few tasks a weak basis for deciding which model will work better in practice.

Read the 115 ms figure with its model version and hardware

The announcement describes a median of approximately 115 ms on an RTX 3090. Its latency chart caption, however, identifies the measured model as v18. That chart cannot establish that the released v19 will run in 115 ms in any environment. Figure 3 in the announcement

The GitHub documentation checked for this article separately reports v19 at a median of 115 ms and a 95th percentile of 299 ms on an RTX 3090 under WSL2. JevBench sends one question per task, so those figures do not include the benefit of sharing a prefix across multiple questions. The announcement's v18 chart and the separate v19 experiment must be read as distinct records.

The Mac figure of 153 ms is conditional too. The developers' detailed record identifies it as the warm median for tasks under 300 tokens, using v19 on an M3 Pro with 36 GB of memory, macOS 26.6, and bf16 through MPS. The first request at those lengths had a median of 310 ms. Across all JevBench tasks, the warm median was 234 ms and the 95th percentile was 2,628 ms. Short-task latency cannot be applied to long documents. Hardware and cold versus warm conditions

Multiple questions about one state can share the computation for that state. They do not read one another's answers, however. Sequential planning that uses an earlier answer to formulate the next question belongs in the application or large LLM. End-to-end savings also depend on loading the model, input length, specialist execution, and how often work returns to the large LLM. How state computation is shared

The CLI and a Python HTTP client use the same request structure

The official entry point is the strands-decider Python package. The following CLI example comes from the announcement, with the command placed on one line for presentation. First use retrieves both the Decider artifacts and the underlying Qwen model. We have not executed it.

pip install strands-decider
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 --state "Help! My payouts have been failing for 3 days! " --choice "Which team should handle this?=billing,sales,retail"

The official output for this request is shown in the probability chart above. Changing the option descriptions or the input can change the selection even for the same topic.

For repeated calls, an application can keep the model loaded with strands-decider serve and send requests to POST /v1/systemone. The official example's Python client calls that HTTP API with the standard library's urllib. Decider is a class in the example's _client.py; this is not an instruction to import that name from the package's top level.

from _client import Decider

decider = Decider("http://127.0.0.1:8099") answers = decider.ask( state="Help! My payouts have been failing for 3 days!", questions={"route": Decider.choice( "Which team should handle this?", {"billing": "payments and invoices", "sales": "pricing and upgrades"}, )}, ) print(answers["route"])

LATENT wrote this unexecuted Python illustration against the official client's API. It assumes that _client.py from a clone's examples/strands/ directory is importable and a Decider server is running at the specified endpoint. For in-process access, the model card also documents StrandsDeciderModel.load(...). Python client / Loading the model

The server binds to 127.0.0.1 by default and has no authentication. The official documentation limits it to local experiments and says behavior under concurrent requests is unverified. GET /health exposes the checkpoint and device, but this is not specified as an OpenAI-compatible chat server. Keeping the model resident does not provide all the infrastructure required for a public, multi-user service. HTTP server scope

Checks before execution pass model scores into application rules

An official example places Decider before a weather tool call. When the user asks about the weather without naming a location, the large LLM is prompted to assume Seattle and call the tool. Decider reads the conversation and proposed call, then assesses whether the arguments are grounded in facts supplied by the user and whether execution is premature without clarification.

The Python handler uses those two numbers to choose Proceed, which allows execution, or Guide, which sends corrective feedback to the model. Strands' intervention API also offers actions including Deny, which blocks execution, and Confirm, which waits for human approval. The handler determines these actions. Decider does not generate the clarification or corrected arguments. Official example implementation / Intervention API

The example's threshold of 0.45 was selected by a person for illustration. Its weather tool returns fixed demo text rather than fetching actual weather. The large LLM in this example calls Amazon Bedrock, so it also requires AWS credentials and access. Decider's classification runs locally; the complete example does not run entirely on the device. Example prerequisites and scope

The official artifacts center on an adapter and scoring weights

The PyPI release checked for this article is 0.1.0 and requires Python 3.10 or later. Inference can target CPU, NVIDIA CUDA, or MPS on Apple silicon. The CUDA instructions assume Linux or WSL2 on Windows. Package release

CPU execution uses reference implementations for some operations, and the documentation describes it as much slower than GPU execution. A 2B model size alone does not imply latency in the 100 ms range on a particular CPU. Execution environments

The checked GitHub revision also has an MLX backend, explicitly selected with --device mlx on Apple silicon. The documentation says that installation extras such as mlx and cuda will ship with the next package release. PyPI 0.1.0 exposes only train and dev; the documented way to try MLX is to install from a source clone. Mixing the published package with instructions for the latest GitHub code can therefore produce mismatched prerequisites.

The official Hugging Face artifacts include a PEFT LoRA adapter, head.safetensors, configuration, and a tokenizer. They do not include the Qwen base weights, which the documentation says require a separate download of about 4.5 GB at first use. That is an approximate file-download size, not the total RAM or VRAM requirement. Working memory also grows with the input and number of questions. Official file list / Loading behavior

The published configuration uses bf16 for the torso, while the scoring head runs in fp32. LoRA is not quantization, and safetensors is not the name of a method for reducing bit width. We found no official GGUF or 4-bit release in the official files. The MLX implementation also rejects merging the adapter into a quantized base model. Published configuration / MLX implementation

A third-party author, fabricant451, has published an experimental Q8_0 GGUF conversion. Its card explicitly requires a custom Strands-aware llama.cpp/wllama runtime and does not assume compatibility with stock GGUF clients. The author's eight reference fixtures are a small conversion check, not a model-quality benchmark. The existence of a GGUF file therefore does not establish that it runs unchanged in ordinary Ollama. LATENT has not downloaded or executed this conversion either. Conversion author's documentation

Code and weight licenses do not cover all training data

Training uses public topic and intent classification data, multi-step tasks from sources such as ContractNLI and MuSiQue, and questions generated and checked by generative models. v19 adds training examples derived from sources including HelpSteer2 to judge whether a response adequately answers a request. HotpotQA is evaluation-only. Source inventory and roles

Synthetic examples and model-produced teacher distributions are committed to the repository. Original public datasets are downloaded and converted separately. Saying that the data is available does not mean that every original source is redistributed in one place under the same terms. The reproduction recipe uses stored synthetic data, so the project says reproducing it requires no paid generation-API calls. Committed and separately downloaded data

The code and released adapter and head are Apache-2.0. The Qwen3.5-2B-Base model is also labeled Apache-2.0. Training sources, however, include a mixture of CC BY, CC BY-SA, and other terms. Some conditions are unrecorded in the official inventory, and some synthetic-data entries list only the generating models' licenses. We do not describe all the data as Apache-2.0. Code license / Released weights and data notes / Base model

We found no link to a standalone Strands Decider 2B paper in the official announcement, model card, or research records checked. The published research evidence consists of architecture documentation, training recipes, and experiment conditions recorded in advance with subsequent results. Cited papers such as HelpSteer2 concern the source datasets; they are not Decider papers. Research record

A narrow decision scope makes the division of work easier to evaluate

A useful starting point is a recurring decision with clear options and errors that can be measured. For extraction-versus-translation routing, an application can compare correct dispatches, returns to the large LLM, end-to-end latency, and mistaken automatic execution. This is LATENT's architectural suggestion, not a measured benefit.

The model can confidently choose the wrong option when the answer is absent, the state lacks a necessary fact, or it misreads the question. The documentation also reports weaknesses with paraphrasing and negation. Long states are shortened to fit the window by default, potentially losing a necessary condition. --strict-window instead allows oversized requests to fail with an error. Question-following limitations / Input-length handling

Adding a fallback option does not guarantee that the model will choose it when needed. Permission boundaries and human approval still need to be enforced by application rules. Decider assists choices within those rules. Unpacking complex requests, asking for missing information, and producing the final explanation remain jobs for the large LLM and the user.

Sources and checked revisions

The official announcement is dated October 1, 2026. We checked GitHub commit 6d5dec6 and the official weight documentation at revision bb282d7. We read documents, configurations, and source text without running the linked code or models. The third-party GGUF account is the conversion author's documentation, separate from official support guarantees.

For a related explanation of software that runs local models, see our TensorFold article.