- AWS's Strands Labs released Strands Decider 2B, a small model that picks from options or scores them, and says it takes a median of around 115 milliseconds on widely available hardware.
- The performance figures are AWS's own, the weather demo was staged, and AWS says the model is significantly worse than reasoning models on complex problems.
- Your team still has to set what score lets an agent proceed, ask a person, or stop, so check which agent actions run unchecked and who approved those thresholds.
The agent that guessed a city
In an AWS demo, a user asks an AI agent for the weather without naming a place. AWS built that agent to be "deliberately eager". So it guesses a city and calls the weather tool anyway.
The guess is staged to show the check working. It is a demonstration of a mechanism, not a measured failure rate for agents.
The mechanism is a small model that reviews the proposed action before the tool runs. The lesson here is that cheap checks move the hard problem. The cost of asking falls. The responsibility for the answer does not.
What Amazon released
AWS's Strands Labs published Strands Decider 2B this week. It is a decision model: it does not write text. It picks from a list of options you give it, or returns a score between 0 and 1.
AWS released the code and training data on GitHub, and the model weights on Hugging Face. VentureBeat reports an Apache 2.0 license. Users supply their own hardware or cloud capacity.
AWS says the release follows TypeSafe AI's launch of a similar model called Jev earlier this month.
How a decision model works
AWS started with an existing language model, Qwen3.5-2B. It removed the part that generates words, called the LM head. In its place it added a small "pointer head" of just over a million parameters. The pointer head scores each answer option.
The base model was adjusted with a rank-16 LoRA adapter. That is a light add-on layer, so the whole model need not be retrained.
The result answers in a single pass. Each answer comes with a reliability score: how sure the model is that the yes or no is right. AWS says frontier language-model APIs do not offer that score. The model can also answer several questions about the same prompt efficiently.
Where it sits in an agent
AWS's example uses the intervention hook in the Strands agent framework. The hook sits in front of each tool call. Its code picks one of four outcomes: let the call go ahead, block it, pause for a person to approve, or return the turn to the agent with advice.
In the weather demo, the model answers two yes-or-no questions. Did the city come from anything the user said? Is it too early to call the tool? The answers send the agent back to ask which city the user meant.
AWS also lists early uses: model routing, tool selection, evaluations, guardrails, memory and policy classification. It describes hybrid agents, where a large model handles hard choices and a decider handles routine ones.
What the numbers say, and what they do not
The performance figures are AWS's own. On JevBench's public set, AWS ranks the model third of 33 in the 2B class for accuracy and calibration. That is first of 30 if models just over 2 billion parameters are excluded.
For speed, AWS says the model makes local decisions in a median of around 115 milliseconds on widely available hardware. That is a general figure. AWS reports no latency for the weather demo, so it is not the time to vet an agent's tool call.
A separate latency chart was measured on a local Nvidia RTX 3090 graphics card against version 18. The released model is version 19. Latency rises roughly in line with task size. On an M3 MacBook, small tasks take a median of about 153 milliseconds.
AWS's text also says "tens of milliseconds" for some questions. The median it gives is higher.
The limits are plain. AWS says the single-pass design makes the model significantly worse than reasoning models on complex problems. It cannot write text, so it does not suit coding, chatbots or summarization.
AWS also calls its weather example "an illustration rather than a recommendation". It says it chose the questions, the cutoff and the follow-up rule manually.
The decision this does not make for you
A confidence score is only useful once someone decides what it triggers. At what score does an agent action proceed, ask a person, or stop? That is a business rule. It is not a model feature.
VentureBeat adds one point on control. Because the model can run on your own infrastructure, each decision need not go to an outside API.
Questions to put to your team
Which actions can our agents take today with no check before they run? Who set the thresholds for proceed, confirm and deny, and who approved them?
Has anyone tested a decision model on our own tasks? A public benchmark score is not a measure of our workload.
When a check says confirm, who receives the request, and how fast do they answer? A cheap check that routes to a slow human has moved the delay, not removed it.
A fast check is not a safe agent. It is a place to put a decision you already had to make.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Amazon Web Services (Strands Labs).





