On October 1, AWS gave away a model that cannot write a sentence. Strands Decider 2B 21 has about 2 billion parameters. For a choice question, you hand it a short list of options and it returns a choice with a calibrated probability for each one. AWS reports a median of around 115 01 milliseconds per decision on an NVIDIA RTX 3090.
On October 2, the ggml.org team added decision-model support to llama.cpp, the engine many builders already use to run models on their own hardware. By October 3, an explainx.ai roundup counted five open decision models it can serve. Two releases in two days turned the small local router into a working server endpoint.
The payoff lands on your API bill. Most agents send every small question to a frontier model: retry or stop, search or ask, tool A or tool B, escalate or answer. Each of those calls costs tokens plus a network round trip. A 2B 04 model on your own machine answers them for the price of electricity.
There are no public production numbers for this pattern yet. The only public score covers 231 02 benchmark tasks, and the model missed 64 of them. The experiment is still cheap enough to run this month.
The Pick-List Principle
The Pick-List Principle sorts every step in an agent loop by the shape of its answer. If the answer fits on a short list you can write in advance, a 2B 04 model picks it. If the answer has to be composed, a frontier model writes it. A third bucket holds the steps no model should own.
Four numbers that size the decide-versus-do split
Pick steps are bounded questions, such as which tool runs next or whether a human is needed. According to its release materials, AWS positions Decider for routing and tool selection. The ggml.org post says llama.cpp scores every option in a single forward pass.
Write steps are open-ended work, like code and anything a customer reads. Frontier prices exist for these calls. For scale, Ant Group's Ling-3.1-flash carries 560 billion 05 total parameters with 25 billion 06 active. The active slice alone is 12.5 25 times the size of all of Decider.
Gate steps, like permissions and spending limits, belong to plain code. A score of 0.91 26 tells you how sure the model is about its pick. It doesn't authorize anything. Authorization still has to come from rules you wrote and tested. Sorting your loop into these buckets is the whole design exercise.
Inside the 115-Millisecond Decision Loop
Think of Decider as a mailroom sorter. It reads the label on each envelope and drops it into one of the slots you built. It never opens the letter. The frontier model works upstairs writing replies, and it bills by the word.
The release materials describe a Qwen3.5-2B torso with a custom pointer head of about 1 million 12 parameters on top. A pointer head is a small output layer that scores each candidate you pass in and points at the best one. Tech Times put the result in its headline: the model "routes without generating any text." No text means no tokens to stream and nothing to parse.
On the llama.cpp side, the new endpoint is /v1/systemone, which follows the System One format tied to TypeSafe's Jev model. You send the current state and the candidate options to score. You get back a probability per option, something like this:
- search_database: 0.91 26
- call_web_search: 0.06
- ask_user: 0.03
One job might trigger 20 27 control decisions. A frontier-only agent pays for 20 or more calls. A split agent could keep 3 to 6 for real work and send the rest to the local sorter. Those splits are illustrations. Your real ratio comes from your own logs.
You can check the money math yourself. Take an illustrative app with 1 million 29 requests a month, and suppose the local layer removes the frontier call on 30% of them. That's 300,000 calls you stop sending.
If each avoided call carried 2,000 input tokens and 500 30 output tokens, you skip 750 million tokens a month. At a placeholder blended rate of $5 31 per million tokens, you get $3,750 a month back. Your provider's rate will differ, so plug yours in. The arithmetic works the same way.
Speed is the second line on the receipt. A remote decision pays for the network and the provider's queue on top of inference. A local one pays mostly for inference.
None of that speed matters if the picks are wrong. Per the release figures, Decider scored 72.3% 03 on JevBench v1, 167 correct out of 231 tasks. That leaves 64 32 misses. Nobody knows yet whether that rate holds on your own tool list, which may look nothing like the benchmark.
The bigger trap is the closed world. If the right action isn't on your list, the model still picks something, and it may sound confident doing it. So the hard thinking shifts to how you write the candidate list and when you escalate. We think the first jobs to hand over are retries and stop checks, where a wrong pick costs one extra call.
Speed, misses and a runtime that moved fast
Control decisions run on hardware you already own.
AWS reports a median of around 115 milliseconds per decision on an RTX 3090. A pointer head scores every option without generating text, so there are no tokens to stream and nothing to parse.
Decider missed 64 of 231 benchmark tasks.
It scored 72.3% 03 on JevBench v1, and nobody knows yet whether that holds on your own tool list. If the right action is missing from the list, the model still picks something and may sound confident doing it.
The category reached a shared runtime in about 17 days 33.
Jev launched around September 16, and by October 3 llama.cpp could serve five open decision models. Once the runtime exists, swapping one decider for another becomes a config change.
2031: Routers Live on Laptops
The dates show how fast this moved. According to The Letter Two, AWS shipped the Strands Agents SDK 16 months 20 before Decider. The Strands labs program followed in February 2026, then a harness on September 21. TypeSafe AI introduced Jev around September 16, per The Register. TechCrunch's October 1 headline described decision models that "flood the web."
Mid-September to October 3 is about 17 days 33. In that window the category went from a fresh launch to five models served by a common open runtime. (A runtime is the software that loads a model and answers requests.) Once the runtime exists, swapping one decider for another is a config change.
The volume math is easy. Savings scale with call volume, and the decider stays the same size. The illustrative app above saves 300,000 29 calls at 1 million requests. At 10 million 34 requests the same 30% rate saves 3 million calls, and the model doesn't get any bigger.
The bet favors the builder. Trying it costs a few days of logging and an Apache 2.0 model on a spare machine. The upside grows with every request you add. Frontier models will keep growing for the Write bucket, as Ling-3.1-flash's 560 billion 05 parameters show, but the Pick bucket has no reason to follow.
Our read for 2031: teams will compare control models on decision accuracy and p99 latency, the way they compare databases today. AWS sells hosted compute, and it just released a router built to run on your Mac. We take that as an early sign the control plane is heading toward commodity status. There's a real counterpoint, though. When deciding and doing are tangled together, a split loop adds handoffs that can fail.
Route Your Retry Decisions Locally First
The build runs in a fixed order you can repeat, called LIST: Log, Install, Shadow, Threshold. First you log, then you install, then you shadow, then you set thresholds. Each step leaves a file you can check.
Step one is the log. For seven days, record every control decision your agent makes. Save the state, the option list, the frontier model's choice and what happened next. You will find decisions you never knew the agent was making.
Step two is the install. Pull the Strands Decider 2B 14 weights from Hugging Face and start llama.cpp's server with them. AWS says it runs on a local CPU, a consumer GPU or an Apple-silicon Mac. Send one logged decision to /v1/systemone and confirm you get a probability per option.
Step three is the shadow run. Shadow mode means the local model answers every live decision while the frontier model keeps driving, and only your logs see the local answer. After a week, count how often the two agree. Read every case where Decider scored above 0.90 35 and still disagreed.
Step four sets the thresholds. Calibration means a 0.90 35 score should be right about 90% of the time, so bucket your logs by score and check. Let Decider act on its own at 0.90 and up, and send anything between 0.60 and 0.90 to the frontier model. Anything under 0.60 goes to review, and every list gets an "escalate" option so the closed-world failure has an exit.
Things will break. A vague tool description will confuse the model, and a long state may get truncated. Start with retries and stop checks, and keep anything that moves money or deletes data in plain code. Rerun the shadow week whenever you add a tool, and count frontier calls avoided each Friday.
Shadow a local decider on your retry calls
- Log every control decision. For seven days, record the state, the option list, the frontier model's choice and what happened next.
- Run Decider in shadow mode. Serve the Strands Decider 2B 35 weights through llama.cpp's /v1/systemone endpoint and let it answer live decisions while the frontier model keeps driving. Read every case where it scored above 0.90 and still disagreed.
- Set thresholds with an exit. Let Decider act alone at 0.90 and up, send 0.60 to 0.90 to the frontier model, and route anything under 0.60 to review. Add an "escalate" option to every list and keep money moves and deletions in plain code.
Give the small model the pick list and keep the pen for the frontier.
Most agent control steps are bounded questions with answers you can write in advance. A 2B 04 model on a spare machine can answer them for the price of electricity, and the savings grow with every request you add. The 64 02 benchmark misses mean the trial belongs in shadow mode, starting with retries and stop checks where a wrong pick costs one extra call. Count the frontier calls you avoid each Friday and let your own logs set the ratio.
