Something strange has happened in AI over the past two weeks, and it's worth the attention of anyone who builds a website or creates content.
A wave of new models launched, every one of them billed as "the strongest" or "number one." But look at what they actually do, and you hit a surprising fact: these hyped new models cannot write a single complete sentence.
They can't chat with you. They can't write an article. They can't answer an open-ended question. They do exactly one thing: take a question you have bounded in advance, pick from a fixed set of options, and tell you how confident it is about each one.
These are called decision models. I want to walk through what's happening, because this isn't just another buzzword. It may be quietly changing the underlying structure of AI systems — including how AI chooses and recommends content.
1. The timeline: a new category exploded in 18 days
The story starts with a model called Jev. On September 15, a company named TypeSafe released it with an intriguing label — a "System One" model.
It generates no text. You give it a piece of "state" (a conversation, an email, a support ticket) and a set of typed questions, and it returns choices, scores, or a yes/no probability. The price is startling: $0.042 per million input tokens, with no charge for output tokens — because it produces no output text at all.
That small model unexpectedly ignited an entire category. What followed was remarkably fast:
- Sept 29: At DevDay, OpenAI quietly introduced the Decisions API (powered by its Luna model), narrowing the capability to "deterministic rulings from a finite set of options." The community promptly nicknamed it a "Jev clone."
- Around the same day, Liquid AI released d1, pitched around "zero output tokens and calibrated probabilities."
- Oct 1: During Birthday Week, Cloudflare released the open-source Clef, along with a faster Clef-flash. Almost simultaneously, Amazon's Strands Labs open-sourced Strands Decider 2B.
- Add a crowd of others, including Fastino Labs' GLiDE, billed as the first "thinking" decision model.
Sources: MarkTechPost, the Cloudflare changelog, and the DevDay 2026 rundown.
In 18 days, one model turned into a "clone war." Even by AI-industry standards, that's rare.
2. Why did "dumber" become the selling point?
You might wonder: aren't large models racing to be smarter, with ever more parameters? Why the sudden competition to be narrower?
The reason is practical. What usually breaks an agent is not an inability to think — it's choosing wrong at some fork in the road.
Consider a customer-support agent. When an email arrives, the first job isn't to compose a fluent reply. It's to answer a string of tiny questions: should this be routed to refunds, technical support, or logistics? Is the customer upset? Does this need to be escalated to a human first?
These questions share two traits: the answers are finite, and they get called at high volume, again and again.
Using a general LLM for this has three problems: it's slow (waiting for word-by-word generation), expensive (you pay for output tokens too), and inconsistent (this time it says "refunds," next time "technical support," and the format may drift).
Decision models target all three. Instead of "writing," they score every option jointly in a single forward pass — think of it as not reading answers aloud, but glancing at them all at once and stamping a probability on each.
The effect is visible in the measured numbers:
- Cloudflare's Clef has a median latency around 209 ms; the lightweight Clef-flash is just 38.8 ms. Jev is around 524 ms. Clef-flash is roughly 13x faster.
- OpenAI's Decisions API hit 76 of 78 steps in one replay, at about 230 ms per decision.
- Amazon's Strands Decider 2B runs locally on an RTX 3090 at a median of about 115 ms.
Tens of milliseconds means these models can sit directly on an agent's "hot path" — before each action, spend a fraction of a second asking "should I do this?", then decide whether the expensive LLM is even needed.
3. What they're genuinely good at — and where they're dumb
Don't be swept along by "number one" and "leading." Let's be even-handed.
The real strengths:
- Speed and cost — the hardest advantages.
- Stable, reusable output. They always return structured results in a fixed shape, not one thing today and another tomorrow, which makes them easy to program against.
- They return probabilities, not just verdicts. If it says "85% refunds," you can set a rule: anything below 80% confidence doesn't get auto-handled — it goes to a human. That confidence signal is genuinely useful in engineering.
- The newer Clef even has a vision encoder that reads screenshots, receipts, and forms, while Jev is currently text-only. Its context window is also 64K, up from Jev's 32K.
But the limits are clear, and they matter:
- They can't reason, can't create, and can't handle a situation that wasn't predefined. If the correct answer simply isn't among your options, it still has to pick one.
- They are not a safety guarantee. A model's nod doesn't make something correct, and researchers note that adversarial text may be able to manipulate the judgment.
- The open-weight ones (Amazon, Cloudflare) need your own hardware to run. They're labeled "free," but you supply the GPU or cloud capacity, and whether self-hosting actually saves money depends on the full bill.
- Amazon itself concedes it places second in accuracy at its size, winning instead on "fully open training recipes." In other words, everyone's "number one" tends to be on a leaderboard they designed.
4. What does this mean for content and GEO?
Now the part you care about. My read: this is still a directional signal, but the direction deserves real attention.
Picture how an AI system may soon be layered. Thousands of small "who should this route to?" decisions are handled by decision models in tens of milliseconds at near-zero cost; only genuinely complex understanding and generation goes to the large model.
And when a decision model makes a "which source should I recommend?" judgment, it is, at its core, scoring and ranking content. Which means:
First, "getting recommended" may increasingly happen inside a layer that writes no text and only computes probabilities. That lines up with the logic behind GEO: you're not chasing a click, you're trying to be the source that scores highest when the AI makes its call.
Second, content a model can pick with high confidence has an edge. What earns a decisive high score? Clear points of view, clean structure, exclusive data, and unambiguous conclusions. Vague, generic "filler content" — correct but empty — loses in probability scoring, because even the model can't tell exactly what you're saying.
Third, there's a striking geopolitical detail. Most of these open-source decision models are built on Alibaba's Qwen: Clef on Qwen3.8-27B, Amazon's Decider on Qwen3.5-2B. A Chinese open-source model is becoming the foundation for the world's AI "decision layer" — a trend worth watching on its own.
A dose of caution, though. These models are still largely English-first, and their performance in other languages — and whether they'll be used directly for public content recommendations — is unsettled. Reworking your whole strategy for them today is premature; ignoring them entirely could mean missing the next inflection point.
5. Three practical takeaways
First, understand the division of labor. AI is moving from "one giant model does everything" toward "cheap small models make countless simple judgments + expensive large models do complex thinking." That's a genuine structural shift, not just marketing.
Second, make content "decisive to judge." Sharp opinions, exclusive information, clear structure, definite conclusions — let a probability-only model hand you a high score without hesitation. This holds regardless of whether you ever touch a decision model.
Third, stay informed without panic or blindly chasing every release. The category is new, the leaderboards will shift, prices will move, and who wins long-term is far from settled.
The curious part is that a race to be "dumber" signals a maturing industry: we're no longer worshipping one all-purpose giant brain, and we're starting to ask seriously which kind of mind fits which kind of work.
For those of us who make content, the answer hasn't really changed. Don't optimize for every new model. Become the source with clear options, solid evidence, and a case that humans — and AI — simply can't avoid choosing.