The news
On October 1, 2026, Amazon Web Services (AWS), part of Amazon (AMZN), and Cloudflare (NET) each released decision models: small AI models that choose among fixed options instead of writing text. Both pitch them for routing and classification steps inside AI agents.
AWS's Strands Labs released Strands Decider 2B, a 2-billion-parameter open-source model built on Qwen3.5-2B, with code on GitHub and weights on Hugging Face. The Strands team says it runs locally with a median latency of about 115 milliseconds on an Nvidia RTX 3090 graphics card and about 153 milliseconds for small tasks on an M3 MacBook. It listed uses including model routing, tool selection, guardrails and policy classification.
SiliconANGLE explained the trade-off: decision models pick from predefined choices and attach a confidence score to each decision, which removes the cost of generating text but means they cannot explain their reasoning. Marc Brooker, an Amazon distinguished engineer, told TechCrunch the models "make a perfect decider for a workflow step."
Cloudflare released Clef, built on a 27-billion-parameter Qwen model, and Clef-flash, built on a 9-billion-parameter one. The Register reported that both are open-weight under the Apache 2.0 license, run on Cloudflare's Workers AI platform or on a customer's own hardware via Hugging Face, and use an API compatible with Jev, the decision model from startup TypeSafe. Cloudflare said median latency across 43 benchmarks was 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash, versus 524.1 milliseconds for Jev.
In one internal test, Cloudflare said Clef took 2.2 seconds to fetch, render and classify a website, while its fastest general large language model, gpt-oss-120b, took 4.7 seconds. The Register noted that Clef's Workers AI price of $0.24 per million tokens is nearly six times Jev's $0.042, and that Cloudflare's benchmark scores are self-reported and have yet to be reproduced for ranking on the official Decision Index.
TypeSafe Chief Executive Diogo Almeida was dismissive of the new entrants, telling TechCrunch that they look more like machine-learning teams trying out an interesting architecture than groups focused on making the technology useful.
The numbers
- Strands Decider size
- 2 billion parameters
- Strands Decider median latency (Nvidia RTX 3090, AWS figure)
- About 115 ms
- Clef-flash median latency (Cloudflare figure)
- 38.8 ms
- Jev median latency (Cloudflare figure)
- 524.1 ms
- Clef price on Workers AI (The Register)
- $0.24 per million tokens
- Jev price (The Register)
- $0.042 per million tokens
Why CEOs should care
For CTOs and heads of engineering, the first step is an inventory. Many agent steps are really multiple-choice questions: which tool to call, which model should handle a request, whether a case needs a human, whether content breaks a policy. Those steps are candidates for a decision model. Test one on your own data, compare its accuracy with the large model doing the job now, and send low-confidence answers back to the larger model or a person.
For CFOs, the savings come from volume. An agent can make many small decisions per task, and each one that moves off a frontier model saves tokens and time. But compare total cost, not list prices: self-hosting Clef needs about 85GB of GPU memory, according to The Register, while Strands Decider is small enough for a laptop. Ask teams for cost per decision and error rate before and after the switch.
CISOs and compliance leads should note the explainability gap. A model that cannot explain its choice needs logging of its inputs, outputs and confidence scores, plus thresholds that trigger human review, especially when it is used as a guardrail. Training data for Clef is not public, The Register reported, which belongs in any model supply-chain review.
The bigger picture
TechCrunch, which called Amazon's model a Jev clone, said dozens of similar models have been produced by researchers since TypeSafe debuted the idea. AWS and Cloudflare entering the category moves it from startup experiment toward a standard part of agent infrastructure: AWS through its open-source Strands agent tooling, Cloudflare through its Workers AI platform.
That puts price pressure on TypeSafe while also validating its idea. For buyers, it adds to the case for a tiered model setup, in which large models handle open-ended reasoning and smaller, cheaper models handle the routine choices around them.
What’s next
Watch for whether the benchmark claims are reproduced for ranking on the official Decision Index, Cloudflare's planned self-serve fine-tuning for Clef, and whether AWS offers Strands Decider as a managed service. Also watch how TypeSafe responds on price and features.
What “Fact-checked” means
Fact-checking means testing a story’s facts against the evidence before it is published. This story went through at least two separate checks before this version was published.
- What we checked
- Its names, figures, dates, job titles, quotes and who said what were checked against the story’s sources, including its main source where it could be opened. The headline was checked for accuracy and overstatement.
- How
- A first check reviewed the whole story. If it passed, a second, skeptical check went back to the sources to look for mistakes in the most important facts. If a check flagged the story, it was edited to fix the problems found, and a separate re-check then reviewed the whole story again.
- Who
- The checks are made by our newsroom, as steps kept separate from the writing, under rules set by our editor, Hussein Mukhtar. A story the checks still flag is held for the editor, who decides whether it is fixed, published or dropped.
- If something is wrong
- “Fact-checked” does not mean error-free. If a material error is found after publication, we correct the story and add a note saying what changed. Report an error
Companies in this story









