Decision models: the small judgments inside useful AI workflows
Why decision models converged so quickly across Jev, OpenAI, Perplexity, and Cloudflare, and where bounded judgments help real workflows.

On this page8 sections
Imagine opening a support queue on Monday morning. One customer cannot reconnect an integration. Another wants an invoice corrected. A third has put both problems in the same message.
A language model could write a thoughtful reply to each of them. Before anyone replies, though, the workflow needs a smaller answer. Which team should handle this? Is information missing? Does the message contain more than one issue?
Those small judgments are where decision models become interesting. They let software ask a question about messy evidence and receive a bounded result it can use directly. The opportunity is to make a specific part of the workflow easier to measure, faster to serve, and easier to compose with ordinary code.
The stories below are fictional design examples. They illustrate where these models might help; they are not reports of customer deployments or tests I have run.
Four providers in sixteen days#
TypeSafe launched Jev on September 15, 2026, describing a model focused on typed, probabilistic decisions and calling the category System One Models. Fourteen days later, on September 29, OpenAI announced a Luna-powered Decisions API in limited preview. That gap is what caught my attention. TypeSafe launch, OpenAI announcement.
By October 1, Perplexity had a documented Decisions API and released weights for pplx-decider-v1-27b. Cloudflare launched Clef and Clef-flash that day, bringing its own decision models to Workers AI and releasing their weights under Apache-2.0. Sixteen days after Jev, the idea had visible offerings from four providers. Perplexity documentation, Perplexity model card, Cloudflare release.
Cloudflare gives the domino-effect framing some direct evidence. Its launch post says the interest in Jev revealed demand for a fast, specific classifier and describes decision experiments it shared in the same week Jev appeared. That establishes an influence on Cloudflare's direction. It does not tell us when OpenAI or Perplexity began their work. Fourteen days between public announcements is not fourteen days to build a model. Cloudflare launch.
The useful signal for builders is the convergence. Several providers now offer a dedicated interface for a familiar job inside software: interpret some evidence and choose among defined outcomes.
What a decision model actually does#
In this article, a decision model means a model or model-backed interface for bounded semantic judgments. The caller supplies evidence, a question, and the possible answers or a rubric. The model evaluates that evidence and returns a typed result. Jev, Perplexity, and Cloudflare expose probabilities as part of that contract. TypeSafe Choice, Perplexity Decisions API, Clef documentation.
For a support router, the options might be billing, technical_support, sales, and needs_review. Their descriptions matter. “Billing owns charges, invoices, and refunds” gives the model a more useful boundary than an unexplained internal queue name.

The application defines both the question and what may happen after the answer.
This overlaps with classification. A typical supervised classifier learns a stable label set from examples. These interfaces let the caller describe questions and candidate answers at request time. That flexibility is useful when several judgments share the same evidence or the decision criteria change with the workflow.
A general LLM can already return a label in valid JSON. OpenAI's Structured Outputs is a baseline worth testing, and its documentation explicitly acknowledges that constrained responses can still contain mistakes. A dedicated decision model needs to earn its place through the trade-off it achieves on your workload. Structured Outputs.
A capable foundation, specialized for a smaller answer#
The shared foundation is the other part I find interesting. Perplexity's pplx-decider-v1-27b and Cloudflare's larger model, Clef, both derive from Qwen3.8-27B. Clef-flash uses Qwen3.5-9B. Cloudflare's Flash option takes a smaller backbone into the same kind of bounded decision interface. Perplexity model card, Clef model card, Clef-flash model card.

Announcement dates show the pace of the category. Backbone relationships show how two providers build on Qwen; neither establishes identical training or performance.
Qwen3.8-27B is a dense vision-language model with image and video understanding and configurable thinking. That is a broad foundation for a product whose answer may be just a distribution over four queues. To me, this is a compelling use of capable open models: specialize their understanding for a job where software needs a small, defined answer. Qwen model card.
Cloudflare makes that specialization concrete. Its model processes the input through the backbone, then a trained schema head reads the resulting representations and scores the options for the supplied questions. The decision interface returns those scores as probabilities instead of generating a free-form response. The useful change reaches beyond formatting JSON; it changes how the answer is produced. Clef architecture.
For a support queue, the larger and Flash options become candidates to compare. Cloudflare reports median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash across 43 benchmark runs. Those are Cloudflare's measurements, not tests I ran or promised response times for your application. Its quality results also vary by task. The smaller model earns its place if it meets the queue's error budget with better latency or cost; the shared Qwen lineage alone cannot answer that. Cloudflare latency report, task results.
Story one: the ticket that belongs to two teams#
Consider Maya, a fictional support lead. Her queue receives this message:
The integration fails after I rotate my API key.
The first candidate use is straightforward. Ask which queue owns the request. The application can accept the recommendation when its evaluation and routing policy allow it, then place the ticket with technical support. A generative model can help draft the response later.
Now change the message:
The integration fails after I rotate my API key, and I think I was charged twice.
A single forced label creates a design problem. Technical support may own the integration issue. Billing may need to investigate the charge. Choosing the most likely queue can hide half the customer's request.

These probabilities are invented for teaching. They are not model outputs or recommended thresholds.
Jev, Perplexity, and Cloudflare document three useful question types. Each fits a different part of Maya's workflow. Perplexity interface, Cloudflare interface.
Choice selects among named options. Ask which queue should receive a ticket, and keep an escape route for cases the other labels do not cover. A review label gives the workflow somewhere to send an unsuitable or ambiguous request. Choice documentation.
Noul assesses a yes/no proposition. For the mixed ticket, ask separately whether it reports an integration problem and whether it raises a billing concern. Both can be true. Code can then coordinate the two queues according to an explicit policy. These questions describe the message; they do not establish whether the account was actually charged twice. That requires trusted transaction records. Noul documentation.
Score evaluates an ordered descriptive rubric. Maya could assess reported impact using levels such as cosmetic issue, degraded feature with a workaround, and blocking issue without a workaround. It returns a distribution over levels and an expected score. Concrete descriptions make the rubric inspectable; an unexplained “severity from one to ten” is harder to validate. Keep the distribution, because an average can conceal ambiguity. Score documentation.
The improvement comes partly from the model and partly from asking a better question. Maya has turned an overloaded decision into smaller judgments the workflow can combine. The actual benefit still needs to be measured through fewer incorrect handoffs and better resolution, rather than label accuracy alone.
Story two: the policy assistant that needs another document#
Now imagine an internal assistant helping Leo prepare for a trip. He asks whether a particular expense is reimbursable. The assistant retrieves a general travel policy, but the exception that governs his case is in another document.
The generative model has enough material to produce fluent prose. It may still lack the evidence needed for this answer.
A decision component could assess a narrower question before drafting: does the supplied evidence explicitly cover the requested expense and the relevant exception? The application can use a missing-evidence signal to retrieve another source or ask Leo for information.
This is a promising use because the judgment is bounded and the next step is reversible. The proposed component checks coverage against specific requirements. It does not certify that a document is authoritative, current, or legally sufficient. Document versions and source identity need separate checks.
The pilot should include passages that look relevant but omit a crucial condition. If the component accepts those passages, it is repeating the original problem in a smaller box. A useful evidence check must earn its value on incomplete and conflicting material.
Story three: the agent that must choose its next tool#
A fictional account assistant receives this request:
Find the latest invoice and tell me whether the payment has arrived.
Several approved read-only tools are available: find invoices, inspect payments, and search the knowledge base. A decision component could select the next tool from that list using the current task state. After the invoice lookup, the application supplies the updated state for the next choice.
This is a local routing decision. A planner may still be needed to organize a longer task or discover an approach. The executor must validate the actual tool arguments, resource, and identity before each call. Choosing the right tool with the wrong account identifier is still a failure.
Change the user's request to “refund that payment,” and the consequences change. A model's selection cannot grant permission to move money. The workflow needs its own authorization and any required approval before executing that action.
The useful architecture combines components. A decision model handles a narrow choice, a generative model explains the result, and software owns state, permissions, and execution.
When the extra component is worth it#
My fit test is five questions:
- Can you name the possible outcomes before making the call?
- Is the relevant evidence available in the input?
- Does the work require interpreting meaning?
- Does it repeat often enough to justify a dedicated component?
- Can you measure mistakes and provide a fallback?
Support triage, document screening, and narrow rubric checks are candidates. A one-off request to write a strategy needs a different kind of model. An exact invoice-total calculation belongs in code. A scheduling problem with hard constraints may need an optimizer.

These components can work together. The diagram describes their role in the application.
Start with the simplest credible baseline. Stable labels and representative training data may favor a conventional classifier. Infrequent judgments may fit comfortably inside an existing general-model call. A dedicated service adds another dependency; lower model charges alone do not establish lower total cost.
There are also documented limits. TypeSafe lists weaknesses in arithmetic, date comparisons, distracting context, adversarial input, and consistency for Jev 1.13. Typed output does not remove those problems. Keep exact checks in code and test messy inputs deliberately. Jev 1.13 limitations.
A probability becomes useful when you test it#
The attractive part of these interfaces is that uncertainty becomes data the application can use. The difficult part is deciding what the numbers justify.
TypeSafe's Choice and Score confidence is derived from the shape of the distribution. It is a different quantity from the largest option probability. Neither field should silently become “the chance this workflow succeeds.” TypeSafe confidence.
Calibration asks whether probabilities correspond to observed frequencies. Across comparable cases assigned a probability near 0.8 to a proposition, about 80% should satisfy that proposition if the predictions are well calibrated. That is a property measured across cases, rather than a guarantee about one ticket. Calibration guide.
For a first pilot, run the support router in shadow mode. Let it recommend routes while people continue handling the tickets. Build the question and tune acceptance rules on separate data from the final test. Include multiple issues, unfamiliar wording, missing evidence, and instructions embedded in the customer message that try to steer the classification.
Compare the decision model with rules and a small general LLM using the same evidence and allowed outcomes. Measure costly routing errors, the share of tickets handled automatically, end-to-end latency, and total cost including review. Report automation coverage alongside error rate. A component that sends everything to a person contributes little automation, even if its accepted predictions look excellent.
Only then does a threshold have meaning. Its job is to express an evaluated trade-off for that task. Revisit it when the model, options, policy, or input population changes.
Maya's goal is a queue that reaches the right people with fewer unnecessary handoffs. A decision model is useful if it improves that workflow. The place to start is one frequent, bounded judgment with an outcome you can observe. Build the rest of the system around it.
I help organizations harness the power of generative AI, with a focus on Microsoft AI and agent frameworks.
Stay in the loop
Continue with your email for subscriber access and future article updates.
Stay in the loop
Continue with your email for subscriber access and future article updates.