K Koda Intelligence
DEEP DIVE DEEP DIVE № 219 · 04 October 2026DOCKODA-20261004-564F38AE135Esha-256 of date + article + 24 checked + 11 computed

Let a 2B model make your agent's small calls

AWS open-sourced Strands Decider 2B 04 on October 1, and on October 2 ggml.org added decision-model support to llama.cpp. By October 3, five open decision models could run on that common runtime. Frontier models keep growing for open-ended work, with Ling-3.1-flash running 25 billion 06 active parameters. We map which agent steps a tiny local router can take over.

6 MIN READ · BY THE KODA EDITORIAL TEAM · TOOLS · AGENT ROUTING
DECIDER SIZE2BVERIFIED CLAIM 21AWS
MEDIAN DECISION115 MSREPORTED CLAIM 01RTX 3090
JEVBENCH V172.3%VERIFIED CLAIM 03RELEASE FIGURES
DECIDER SIZE2BAWS MEDIAN DECISION115 MSRTX 3090 JEVBENCH V172.3%RELEASE FIGURES TASKS MISSED64OF 231 POINTER HEAD1MPARAMETERS LLAMA.CPP SUPPORTOCT 2GGML.ORG CATEGORY WINDOW17 DAYSMID-SEPT TO OCT 3

On October 1, AWS gave away a model that cannot write a sentence. Strands Decider 2B 21 has about 2 billion parameters. For a choice question, you hand it a short list of options and it returns a choice with a calibrated probability for each one. AWS reports a median of around 115 01 milliseconds per decision on an NVIDIA RTX 3090.

On October 2, the ggml.org team added decision-model support to llama.cpp, the engine many builders already use to run models on their own hardware. By October 3, an explainx.ai roundup counted five open decision models it can serve. Two releases in two days turned the small local router into a working server endpoint.

The payoff lands on your API bill. Most agents send every small question to a frontier model: retry or stop, search or ask, tool A or tool B, escalate or answer. Each of those calls costs tokens plus a network round trip. A 2B 04 model on your own machine answers them for the price of electricity.

There are no public production numbers for this pattern yet. The only public score covers 231 02 benchmark tasks, and the model missed 64 of them. The experiment is still cheap enough to run this month.

The Pick-List Principle

The Pick-List Principle sorts every step in an agent loop by the shape of its answer. If the answer fits on a short list you can write in advance, a 2B 04 model picks it. If the answer has to be composed, a frontier model writes it. A third bucket holds the steps no model should own.

LOCAL ROUTER MATH · OCTOBER 2026AWS · GGML.ORG · KODA ILLUSTRATIONBASE: 24 CHECKED + 11 COMPUTED, 4 SHOWN

Four numbers that size the decide-versus-do split

Median decision latency AWS · NVIDIA RTX 3090 REPORTED CLAIM 01
115 ms
JevBench v1 accuracy Release figures · 167 of 231 tasks VERIFIED CLAIM 03
72.3%
Frontier calls avoided Koda illustration · 30% of 1M requests COMPUTED CLAIM 29
300,000
Monthly token spend kept Koda illustration · $5 per million placeholder COMPUTED CLAIM 31
$3,750

Pick steps are bounded questions, such as which tool runs next or whether a human is needed. According to its release materials, AWS positions Decider for routing and tool selection. The ggml.org post says llama.cpp scores every option in a single forward pass.

Write steps are open-ended work, like code and anything a customer reads. Frontier prices exist for these calls. For scale, Ant Group's Ling-3.1-flash carries 560 billion 05 total parameters with 25 billion 06 active. The active slice alone is 12.5 25 times the size of all of Decider.

Gate steps, like permissions and spending limits, belong to plain code. A score of 0.91 26 tells you how sure the model is about its pick. It doesn't authorize anything. Authorization still has to come from rules you wrote and tested. Sorting your loop into these buckets is the whole design exercise.

Inside the 115-Millisecond Decision Loop

Think of Decider as a mailroom sorter. It reads the label on each envelope and drops it into one of the slots you built. It never opens the letter. The frontier model works upstairs writing replies, and it bills by the word.

A score of 0.91 tells you how sure the model is about its pick. It doesn't authorize anything. Authorization still has to come from rules you wrote and tested.· KODA EDITORIAL TEAM · OCTOBER 2026

The release materials describe a Qwen3.5-2B torso with a custom pointer head of about 1 million 12 parameters on top. A pointer head is a small output layer that scores each candidate you pass in and points at the best one. Tech Times put the result in its headline: the model "routes without generating any text." No text means no tokens to stream and nothing to parse.

On the llama.cpp side, the new endpoint is /v1/systemone, which follows the System One format tied to TypeSafe's Jev model. You send the current state and the candidate options to score. You get back a probability per option, something like this:

  • search_database: 0.91 26
  • call_web_search: 0.06
  • ask_user: 0.03

One job might trigger 20 27 control decisions. A frontier-only agent pays for 20 or more calls. A split agent could keep 3 to 6 for real work and send the rest to the local sorter. Those splits are illustrations. Your real ratio comes from your own logs.

You can check the money math yourself. Take an illustrative app with 1 million 29 requests a month, and suppose the local layer removes the frontier call on 30% of them. That's 300,000 calls you stop sending.

If each avoided call carried 2,000 input tokens and 500 30 output tokens, you skip 750 million tokens a month. At a placeholder blended rate of $5 31 per million tokens, you get $3,750 a month back. Your provider's rate will differ, so plug yours in. The arithmetic works the same way.

Speed is the second line on the receipt. A remote decision pays for the network and the provider's queue on top of inference. A local one pays mostly for inference.

None of that speed matters if the picks are wrong. Per the release figures, Decider scored 72.3% 03 on JevBench v1, 167 correct out of 231 tasks. That leaves 64 32 misses. Nobody knows yet whether that rate holds on your own tool list, which may look nothing like the benchmark.

The bigger trap is the closed world. If the right action isn't on your list, the model still picks something, and it may sound confident doing it. So the hard thinking shifts to how you write the candidate list and when you escalate. We think the first jobs to hand over are retries and stop checks, where a wrong pick costs one extra call.

Speed, misses and a runtime that moved fast

LOCAL SPEED
115 01ms

Control decisions run on hardware you already own.

AWS reports a median of around 115 milliseconds per decision on an RTX 3090. A pointer head scores every option without generating text, so there are no tokens to stream and nothing to parse.

CLOSED WORLD
64 02

Decider missed 64 of 231 benchmark tasks.

It scored 72.3% 03 on JevBench v1, and nobody knows yet whether that holds on your own tool list. If the right action is missing from the list, the model still picks something and may sound confident doing it.

RUNTIME RACE
17

The category reached a shared runtime in about 17 days 33.

Jev launched around September 16, and by October 3 llama.cpp could serve five open decision models. Once the runtime exists, swapping one decider for another becomes a config change.

2031: Routers Live on Laptops

The dates show how fast this moved. According to The Letter Two, AWS shipped the Strands Agents SDK 16 months 20 before Decider. The Strands labs program followed in February 2026, then a harness on September 21. TypeSafe AI introduced Jev around September 16, per The Register. TechCrunch's October 1 headline described decision models that "flood the web."

Mid-September to October 3 is about 17 days 33. In that window the category went from a fresh launch to five models served by a common open runtime. (A runtime is the software that loads a model and answers requests.) Once the runtime exists, swapping one decider for another is a config change.

The volume math is easy. Savings scale with call volume, and the decider stays the same size. The illustrative app above saves 300,000 29 calls at 1 million requests. At 10 million 34 requests the same 30% rate saves 3 million calls, and the model doesn't get any bigger.

The bet favors the builder. Trying it costs a few days of logging and an Apache 2.0 model on a spare machine. The upside grows with every request you add. Frontier models will keep growing for the Write bucket, as Ling-3.1-flash's 560 billion 05 parameters show, but the Pick bucket has no reason to follow.

Our read for 2031: teams will compare control models on decision accuracy and p99 latency, the way they compare databases today. AWS sells hosted compute, and it just released a router built to run on your Mac. We take that as an early sign the control plane is heading toward commodity status. There's a real counterpoint, though. When deciding and doing are tangled together, a split loop adds handoffs that can fail.

Route Your Retry Decisions Locally First

The build runs in a fixed order you can repeat, called LIST: Log, Install, Shadow, Threshold. First you log, then you install, then you shadow, then you set thresholds. Each step leaves a file you can check.

Step one is the log. For seven days, record every control decision your agent makes. Save the state, the option list, the frontier model's choice and what happened next. You will find decisions you never knew the agent was making.

Step two is the install. Pull the Strands Decider 2B 14 weights from Hugging Face and start llama.cpp's server with them. AWS says it runs on a local CPU, a consumer GPU or an Apple-silicon Mac. Send one logged decision to /v1/systemone and confirm you get a probability per option.

Step three is the shadow run. Shadow mode means the local model answers every live decision while the frontier model keeps driving, and only your logs see the local answer. After a week, count how often the two agree. Read every case where Decider scored above 0.90 35 and still disagreed.

Step four sets the thresholds. Calibration means a 0.90 35 score should be right about 90% of the time, so bucket your logs by score and check. Let Decider act on its own at 0.90 and up, and send anything between 0.60 and 0.90 to the frontier model. Anything under 0.60 goes to review, and every list gets an "escalate" option so the closed-world failure has an exit.

Things will break. A vague tool description will confuse the model, and a long state may get truncated. Start with retries and stop checks, and keep anything that moves money or deletes data in plain code. Rerun the shadow week whenever you add a tool, and count frontier calls avoided each Friday.

DOJO · BUILD THIS WEEKEND

Shadow a local decider on your retry calls

  1. Log every control decision. For seven days, record the state, the option list, the frontier model's choice and what happened next.
  2. Run Decider in shadow mode. Serve the Strands Decider 2B 35 weights through llama.cpp's /v1/systemone endpoint and let it answer live decisions while the frontier model keeps driving. Read every case where it scored above 0.90 and still disagreed.
  3. Set thresholds with an exit. Let Decider act alone at 0.90 and up, send 0.60 to 0.90 to the frontier model, and route anything under 0.60 to review. Add an "escalate" option to every list and keep money moves and deletions in plain code.
Practice: Design an AI-Assisted Workflow
THE BOTTOM LINE

Give the small model the pick list and keep the pen for the frontier.

Most agent control steps are bounded questions with answers you can write in advance. A 2B 04 model on a spare machine can answer them for the price of electricity, and the savings grow with every request you add. The 64 02 benchmark misses mean the trial belongs in shadow mode, starting with retries and stop checks where a wrong pick costs one extra call. Count the frontier calls you avoid each Friday and let your own logs set the ratio.

LISTEN · AUDIO BRIEFINGThe conversation · ~18 min
WATCH · VISUAL NARRATIVEAnimated breakdown · ~6 min
PLAY · YOUTUBE
EDITORIAL RECEIPTKODA-20261004-564F38AE135E
As of04 October 2026MethodClaim extraction, dated-evidence review, and temporal consistency gate.CorrectionsContact the Koda desk
EVIDENCE24 CHECKED + 11 COMPUTED · 9 VERIFIED · 14 REPORTED · 1 FAILED
9 verified14 reported1 failed11 computed
  1. 01AWS reports a median of around 115 milliseconds per decision on an NVIDIA RTX 3090.REPORTEDMOSTLY TRUEBENCHMARKCORRECTED IN COPYdev.to
  2. 02The only public score for Strands Decider 2B covers 231 benchmark tasks, and the model missed 64 of themREPORTEDMOSTLY TRUEBENCHMARKstrandsagents.com
  3. 03Per release figures, Strands Decider 2B scored 72.3% on JevBench v1, with 167 correct out of 231 tasksVERIFIEDTRUEBENCHMARKtheaieconomy.substack.com
  4. 04On October 1, AWS released the Strands Decider 2B model for freeREPORTEDMOSTLY TRUEMODELdev.to
  5. 05Ant Group's Ling-3.1-flash has 560 billion total parametersVERIFIEDTRUEMODELtechnode.com
  6. 06Ant Group's Ling-3.1-flash has 25 billion active parametersREPORTEDMOSTLY TRUEMODELmindstudio.ai
  7. 07According to release materials, Strands Decider 2B is built on a Qwen3.5-2B torsoREPORTEDMOSTLY TRUEMODELtheaieconomy.substack.com
  8. 08Strands Decider 2B cannot generate textVERIFIEDTRUEFEATUREstrandsagents.com
  9. 09For a choice question, you hand it a short list of options and it returns a choice with a calibrated probability for each one.REPORTEDMOSTLY TRUEFEATURECORRECTED IN COPYstrandsagents.com
  10. 10On October 2, the ggml.org team added decision-model support to llama.cppREPORTEDMOSTLY TRUEFEATUREhuggingface.co
  11. 11llama.cpp is an engine many builders use to run models on their own hardwareVERIFIEDTRUEFEATUREfreedom.tech
  12. 12According to release materials, Strands Decider 2B uses a custom pointer head of about 1 million parametersVERIFIEDTRUEFEATUREtheaieconomy.substack.com
  13. 13Strands Decider 2B is released under the Apache 2.0 licenseVERIFIEDTRUEFEATUREstrandsagents.com
  14. 14Strands Decider 2B weights are available on Hugging FaceVERIFIEDTRUEFEATUREhuggingface.co
  15. 15By October 3, an explainx.ai roundup counted five open decision models that llama.cpp can serveREPORTEDMOSTLY TRUEATTRIBUTIONexplainx.ai
  16. 16According to AWS release materials, AWS positions Strands Decider 2B for routing, tool selection, policy checks and memory decisionsREPORTEDMOSTLY TRUEATTRIBUTIONgithub.com
  17. 17A Tech Times headline said Strands Decider 2B "routes without generating any text"REPORTEDMOSTLY TRUEATTRIBUTIONtechtimes.com
  18. 18Claim removed during the check; its text is not republished.FAILEDMOSTLY FALSEATTRIBUTIONCUT FROM COPYbuilder.aws.com
  19. 19AWS says Strands Decider 2B runs on a local CPU, a consumer GPU or an Apple-silicon MacREPORTEDMOSTLY TRUEATTRIBUTIONbuilder.aws.com
  20. 20According to The Letter Two, AWS shipped the Strands Agents SDK 16 months before Strands Decider 2BREPORTEDMOSTLY TRUEHISTORYaws.amazon.com
  21. 21Strands Decider 2B has about 2 billion parametersVERIFIEDTRUESTATstrandsagents.com
  22. 22There are no public production numbers yet for using a small local decision model to route agent decisionsREPORTEDMOSTLY TRUESTATaiagentstore.ai
  23. 23According to a ggml.org post, llama.cpp scores every option of a decision model in a single forward passREPORTEDMOSTLY TRUEFEATUREhuggingface.co
  24. 24llama.cpp's new decision-model endpoint is /v1/systemoneVERIFIEDTRUEFEATUREhuggingface.co
  25. 25Ling-3.1-flash's 25 billion active parameters are 12.5 times the roughly 2 billion parameters of Strands Decider 2BCOMPUTEDCOMPUTED
  26. 26Illustrative example output from the /v1/systemone endpoint: search_database 0.91, call_web_search 0.06, ask_user 0.03COMPUTEDCOMPUTED
  27. 27Illustrative example: one agent job might trigger 20 control decisions, so a frontier-only agent pays for 20 or more callsCOMPUTEDCOMPUTED
  28. 28Illustrative example: a split agent could keep 3 to 6 of 20 calls on a frontier model and route the rest to a local decision modelCOMPUTEDCOMPUTED
  29. 29Illustrative example: an app with 1 million requests a month that removes the frontier call on 30% of them avoids 300,000 callsCOMPUTEDCOMPUTED
  30. 30Illustrative example: 300,000 avoided calls at 2,000 input tokens and 500 output tokens each skips 750 million tokens a monthCOMPUTEDCOMPUTED
  31. 31Illustrative example: 750 million tokens at a placeholder blended rate of $5 per million tokens saves $3,750 a monthCOMPUTEDCOMPUTED
  32. 32Strands Decider 2B's 167 correct out of 231 JevBench v1 tasks leaves 64 missesCOMPUTEDCOMPUTED
  33. 33The period from mid-September to October 3 is about 17 daysCOMPUTEDCOMPUTED
  34. 34Illustrative example: at 10 million requests, a 30% routing rate saves 3 million frontier callsCOMPUTEDCOMPUTED
  35. 35Article's suggested thresholds: let Strands Decider 2B act alone at scores of 0.90 and up, send 0.60 to 0.90 to the frontier model, and send under 0.60 to reviewCOMPUTEDCOMPUTED

Every claim listed here was extracted from this article and checked against live sources before publication. The verdict is the checker's, not the writer's. Claims the check removed are counted but not republished.

Audit receipt KODA-20261004-564F38AE135E
Filed underToolsDeep Dive04 October 2026
Browse the Deep Dive archive

Get the morning Signal

191 editions so far, one a day. Unsubscribe anytime.