TL;DR
-
A decision model answers questions; it does not write. You send a state (text, JSON, sometimes an image) and typed questions. It returns one bounded decision signal per question with probability information over the possible outcomes, without autoregressive token generation.
-
Three question types cover most decisions.
choicepicks one label from a list you supply,scoreplaces the state on an ordered rubric, andnoulreturns the probability that a yes/no statement is true. Options are defined per request, so no retraining is needed for new labels. The yes/no type has different names across providers:noul(TypeSafe and compatible servers),predicate(OpenAI), andboolean(Vercel AI SDK). -
The difference from structured outputs is mechanical, not cosmetic. An LLM with structured outputs generates a JSON answer token by token. A decision model scores every option you listed, usually with a prefill-only pass and no decode loop, so it cannot return a label you did not offer and it gives you a full distribution to threshold.
-
Probability, confidence, and correctness are three different numbers. The distribution is the model's probability for each option. A
confidencefield, where one exists, is an implementation-specific summary of how concentrated that distribution is. Neither means "probability this answer is right" until you calibrate it on your own labeled data. -
A de facto compatibility contract has formed around TypeSafe's API. TypeSafe introduced
POST /v1/systemonewith Jev in September 2026. Within three weeks Cloudflare (Clef), open projects (Kev, Laya), llama.cpp, Ollama, and Pydantic AI supported that request shape, and SGLang shipped a compatible route next to its own/v1/decisions. OpenAI's separate Decisions API (POST /v1/decisions) entered public beta on 2026-10-06, and Vercel's AI Gateway serves both shapes. Limits, response fields, and usage accounting still differ. -
Use it for routing, triage, gating, and checks, not for anything that must be written. Pair it with a language model: the decision model handles bounded decisions that clear a calibrated threshold, and the LLM or a human takes the uncertain and open-ended cases.
On this page
- Why This Matters
- The Problem Decision Models Solve
- How We Got Here
- What Is a Decision Model?
- How Decision Models Work
- Architecture
- Step-by-Step Flow
- Real Production Example
- Design Decisions
- Comparisons
- Common Mistakes
- Where It Breaks Down
- When NOT to Use Decision Models
- Running in Production
- Related Guides
- Interview Questions
- Key Takeaways
- FAQs
- References
- Further Reading
Why This Matters
Most calls an AI system makes are not requests for prose. They are small decisions: which team owns this ticket, is this message urgent, does this tool call stay inside the user's request, is this document a contract, should the agent retry or stop. In a typical agent loop, those decisions outnumber the steps that actually need generated text.
Teams usually make these decisions with a general-purpose LLM and a schema. That works, but it carries costs that compound at volume:
- Latency on the hot path. An LLM call that reasons and then emits JSON takes seconds. A router or gate that runs on every request makes the whole product slower.
- Spend on throwaway tokens. Reasoning tokens and JSON syntax are billed even though the application keeps a single label.
- No usable uncertainty signal. A generated
"urgent": truedoes not say how sure the model was. Engineers bolt on self-reported confidence scores, which are often poorly calibrated, or run the same prompt several times and vote. - Shape failures. Even with constrained decoding, providers differ in support, and prompt-only JSON still fails sometimes.
Decision models target exactly this slice of work. The call returns a distribution over the options you defined, so code can branch on it directly and send the uncertain tail to a stronger model or a person. When the decision model is right often enough above your threshold, much of the traffic never reaches the expensive path.
This matters even if you never deploy one. The pattern behind it, separating bounded decisions from open-ended generation and gating actions on calibrated probability, is a sound way to structure any agent or workflow.
The Problem Decision Models Solve
Consider a support pipeline that must route each inbound message, flag urgency, and score customer frustration before an agent drafts a reply.
With a general LLM, you write a prompt, attach a JSON Schema, and parse the result. You get one label per field. If you want to know whether the model was unsure, you either ask it to rate itself (unreliable) or call it several times (slow and expensive). If the provider's strict mode does not cover a schema feature you use, you add retries.
With a trained classifier, you get fast class probabilities or scores (calibration may take extra work, such as temperature scaling), but only for the labels it was trained on. Adding a "security" team means collecting labeled data and retraining. Each new decision (urgency, frustration, refund intent) is a separate model to build and maintain.
With zero-shot classifiers (NLI models or embedding similarity), you can supply labels at request time, but accuracy drops on nuanced instructions, ordinal rubrics are awkward, and you still build one pipeline per question.
A decision model sits between these:
| Need | General LLM + schema | Trained classifier | Decision model |
|---|---|---|---|
| Labels defined at request time | Yes | No | Yes |
| Natural-language instructions per field | Yes | No | Yes |
| Probability per option | No (not natively) | Class probabilities or scores | Yes |
| Can return a label you did not offer | Possible without strict | No | No |
| Answer generated token by token | Yes | No | No |
| Several questions about one input | One call, one JSON | One model each | One request, all of them |
| Open-ended text (summary, reply, reason) | Yes | No | No — use an LLM for this |
Neither column's probabilities are guaranteed to be calibrated on your traffic; that is something you measure, whichever model produces them.
The problem it solves is narrow but common: fast, cheap answers with a probability per option, for bounded questions whose options change faster than you can retrain a classifier.
How We Got Here
Diagram: Evolution of typed decisions with models
flowchart TB
A["Pre-2019<br/>Classifier per label set"]
B["2019<br/>Zero-shot classification<br/>via NLI entailment"]
C["2023–2024<br/>LLM JSON mode and<br/>schema-constrained output"]
D["2025 research<br/>Labels set at request time;<br/>confidence-based deferral"]
E["Sep 15, 2026<br/>TypeSafe ships Jev and<br/>POST /v1/systemone"]
F["Sep 29 – Oct 2, 2026<br/>OpenAI Decisions preview;<br/>Kev 1.0, Clef, Strands Decider,<br/>SGLang, Ollama support"]
G["Oct 5–6, 2026<br/>llama.cpp v0.6.0;<br/>OpenAI Decisions API<br/>public beta"]
A --> B --> C --> D --> E --> F --> G
The recent development is not classification itself, but a model and API pattern built for request-time typed decisions, per-option probabilities, and non-autoregressive inference.
Classification is one of the oldest jobs in machine learning. Through the 2010s, teams trained a model per label set: spam filters, intent classifiers, sentiment models. They were fast and cheap, though modern neural classifiers often needed recalibration before their scores could be read as probabilities (Guo et al., 2017), and they were rigid.
Around 2019, researchers showed that natural language inference (NLI) models could classify text zero-shot by testing whether "This text is about {label}" is entailed. Embedding similarity offered another route. Both let engineers define labels at request time, at some cost in accuracy.
Large language models then absorbed most of this work. With JSON mode and later schema-constrained decoding, an LLM could classify, extract, and score in one call with rich instructions. The trade-off was latency, cost, and the absence of a native probability per option.
By 2025, two research threads had made the ingredients familiar without defining the category. Schema-driven encoders such as GLiNER2 (EMNLP 2025) classified text against labels supplied at inference time, several tasks per call, on a CPU-sized model. Work on cascades and deferral, such as Cascaded Language Models for Cost-Effective Human–AI Decision-Making, used calibrated confidence to decide whether a small model answers, a larger model takes over, or a person reviews. Convai Innovations, which later published Laya, released confidence-aware routing research the same year. None of this work proposed today's decision-model API.
In September 2026, TypeSafe AI released Jev and called it the first "System One model", after Daniel Kahneman's fast, intuitive System 1. TypeSafe described a model that "gives up string generation" in exchange for speed and typed answers with probabilities, served over POST /v1/systemone.
Note
Terminology. This guide uses decision model as the generic engineering term, as Cloudflare, ggml-org (llama.cpp), and Pydantic AI do. System One is TypeSafe's name for its model class, and
/v1/systemoneis TypeSafe's API, which other projects implement for compatibility. It is not an industry standard.
Within weeks:
- Jared Palmer released Kev 1.0, an open-source family from 0.8B to 27B parameters that implements the same API. The smaller models are LoRA adapters plus a pointer head on frozen Qwen 3.5 bases; Kev-27B fine-tunes every weight of Qwen3.8-27B.
- Convai Innovations published Laya, an encoder-based decision model built on ModernBERT that answers the same three question types.
- Cloudflare released Clef and Clef-flash on Workers AI with open weights under Apache 2.0. Cloudflare says Clef follows the System One API, so a Jev integration can switch by changing the endpoint and model.
- AWS's Strands Labs released Strands Decider 2B, an open-source model that replaces the language-model head of Qwen3.5-2B with a small pointer head and fine-tunes the backbone with a LoRA adapter. The weights, training data, and scripts are published.
- Liquid AI released open-weight d1-3B (2026-10-07), built on LFM2.5-VL-3B, which answers typed questions over text and images with a 32K context in one forward pass and zero output tokens. Liquid reports a Decision Index 0.2.1 score of 48.57, the best among models under 10B parameters, and latency of 16 ms on Jetson AGX Thor and 50 ms on Jetson Orin Nano. An experimental d1-omni-600M, built on LFM2.5-Encoder-350M, accepts text with an image or text with audio. llama.cpp supported both on day one. The weights use the LFM Open License v1.0, which allows free commercial use only for organizations under $10M in annual revenue, so check the terms before production use.
- llama.cpp shipped a
/v1/systemoneserver endpoint in its v0.6.0 release (2026-10-05), with Clef (text and vision) and Nimble among the supported models. SGLang v0.5.21 shipped/v1/decisions, which turns an existing LLM or VLM into a classifier and scorer, plus a System One–compatible route. Ollama added decision models over/v1/systemone, and Pydantic AI addedDecisionModelandSystemOneModelclasses. - Vercel added an experimental
experimental_decidefunction to the AI SDK (ai7.0.128 or later) and decision models to AI Gateway, which also exposes an OpenAI-compatible/decisionsendpoint. - OpenAI previewed a Decisions API at DevDay and released it in public beta on 2026-10-06. It uses a dedicated
POST /v1/decisionsendpoint, supports onlygpt-6-lunafor now, and bills input tokens only. OpenAI says it expects general availability in the coming weeks.
That pace is the reason for caution as much as interest. The engineering idea is solid and multiply implemented. The benchmarks, naming, and long-term API shape are still settling.
What Is a Decision Model?
A decision model is a model that takes a state and a set of typed questions, and returns a bounded decision signal for each question, with probability information over the possible outcomes. It does not generate free text. A choice answer selects one option from the supplied set, a noul answer is P(true) for the supplied yes/no statement, and a score answer is a probability-weighted value derived from the supplied ordered levels, so it can fall between them.
Inputs
| Input | What it is |
|---|---|
| State | The material to judge: a string, or a JSON object or array (chat logs, records, app state). Some models also accept images. |
| Questions | A map of named questions. You choose the keys; answers come back under the same keys. |
| Instructions | Natural-language wording for each question ("Is this support request urgent?"). |
| Criteria | The options: labels with optional descriptions, ordered rubric levels, or true/false wording. |
The three question types
TypeSafe's /v1/systemone API defines three primitives, and the compatible implementations (Laya, Ollama, llama.cpp, Clef, Kev) use the same three:
| Type | Asks | Criteria shape | Main answer fields |
|---|---|---|---|
choice |
Which one of these options? | Map of label → description (or null), or a label list |
choice (argmax label), probabilities, usually confidence* |
score |
Where on this ordered rubric? | Ordered list of level descriptions, lowest first | score (probability-weighted mean level), probabilities, legend, confidence* |
noul |
Is this statement true? | Optional { "true": "...", "false": "..." } wording |
noul — the model's probability that the statement is true |
* confidence is a summary of the distribution whose formula differs by implementation; it is not the probability that the answer is correct. Most implementations return no confidence for noul.
OpenAI's Decisions API uses the same three ideas with different names and fields: predicate returns probability (the estimate that the condition is true), while choice and score return probabilities and confidence. Images must be sent as inline base64 data URLs. The Vercel AI SDK calls the yes/no type boolean. Map field names explicitly when you switch providers.
A choice question generalizes multi-class classification with labels supplied at request time. A score question handles ordinal judgments (severity, frustration, quality) where "Major" is closer to "Critical" than to "No impact". A noul question returns a single probability for a yes/no statement; how well that probability is calibrated depends on the model and on your data.
Outputs
The response contains an answer per question and a usage block. There is no autoregressive decode, so nothing in the answer is generated text. Usage accounting is provider-specific: llama.cpp and Laya report output_tokens: 0, while TypeSafe's API reference shows small non-zero counts and Ollama documents the field as tokens produced internally for scoring. Read each provider's reference before parsing fields or estimating cost.
Three numbers in a response are easy to conflate:
- Probability is the distribution the model assigns over the options:
probabilitiesforchoiceandscore, andnoulas P(true). It describes what the model believes, not whether the model is right. - Confidence is an implementation-specific summary of how concentrated that distribution is. TypeSafe and Kev rescale the top probability (or the spread around the mean level) against a uniform answer; Ollama and Laya use normalized entropy for
choice; Pydantic AI reports a margin from its decision threshold. Values from different backends are not comparable, and none of them is P(correct). - Calibrated correctness probability is a signal you validated against labeled outcomes on your target traffic, for example "of answers with top probability at or above 0.9, 97% were right". No API returns this by default; you build it by measuring one of the first two signals against your own labels.
You can threshold top probability, P(true), a provider's confidence, or another explicitly defined signal, but none of them is a calibrated probability by default. Any signal used for automation must be validated against labeled outcomes on your target workload.
What a decision model is not
- Not a chat model. It cannot write a reply, summary, or explanation.
- Not an extractor of arbitrary values. It cannot return a free-form string, date, or unbounded number. Those still need an LLM with structured outputs.
- Not a policy engine. It estimates probabilities. Your code decides what they mean and what to do with them.
How Decision Models Work
Scoring instead of generating
An autoregressive LLM produces an answer one token at a time, each step conditioned on the previous ones. Even with constrained decoding, it is sampling a path through the space of valid outputs.
A decision model instead treats each question's options as a closed set and scores all of them. Implementations published so far follow two broad designs.
Decoder backbone, prefill only, plus a scoring head. A causal language model reads the state, question, and options in prefill only, with no decode loop. A small head then scores each option against the question:
- Kev's README describes a Qwen backbone run prefill-only, LoRA adapters on the smaller checkpoints (Kev-27B is fully fine-tuned), and a pointer head that scores each option's closing token against the question's final token. A softmax gives the distribution.
- Cloudflare describes Clef as a frozen Qwen backbone (Qwen3.8-27B for Clef, Qwen3.5-9B for Clef-flash) with a routing head and low-rank adapters. Each valid option gathers evidence from the prompt, fields cross-attend to one another, and the decision step is non-autoregressive.
Encoder backbone with option markers. Laya places a [MASK] marker per option, reads the whole sequence bidirectionally with a ModernBERT encoder, refines it with a small decision head, and reads one logit per marker. A learned type embedding tells the model whether it is answering choice, score, or noul. This design is far smaller (hundreds of millions of parameters) with a shorter context window.
In both designs, the model cannot produce an option you did not list, because there is no generation step to produce one.
Several questions travel in one request, but how they are executed varies. TypeSafe evaluates questions "in parallel and in isolation"; llama.cpp answers them independently, with some models reading the state only once; Laya batches all questions into one forward pass; current Kev models run each question as its own sequence. Clef is the notable exception: Cloudflare says its fields cross-attend, so answers can inform one another. Unless your provider documents otherwise, assume questions are independent and one answer cannot see another.
Calibration
A probability is only useful for automation if it means what it says: of all answers given at 0.9, about 90% should be right. Several decision-model builders report treating calibration as a training target:
- Proper scoring rules. Clef's post-training pairs label-smoothed cross-entropy with a Brier loss. Laya's specialist checkpoint uses a reward built from strictly proper scoring rules, which are maximized only by honest probabilities.
- Temperature scaling. Kev fits a temperature per checkpoint on held-out data and ships it as the default. Its README notes that a single temperature cannot change which answers are ranked as more or less certain.
- Reinforcement Learning for Calibrated Decisions (RLCD). TypeSafe introduced the term for its training method; Cloudflare and Laya describe their own RLCD variants that reward calibrated distributions and give partial credit to adjacent ordinal levels.
The results are uneven. Laya's model card reports an expected calibration error of 0.213 for its base checkpoint and advises treating its confidence as uncalibrated, and a published 300-example pilot found Jev worse calibrated than a GLiNER2.5 baseline on one emotion dataset. Calibration measured on a vendor's data does not transfer automatically to yours. Re-measure on your own labeled set before setting thresholds.
Diagram: One decision-model request
sequenceDiagram
participant App
participant DM as Decision model
participant Head as Scoring head
App->>DM: state + typed questions
DM->>DM: Prefill only, no decoding
DM->>Head: Option representations
Head-->>DM: Logit per option, softmax
DM-->>App: Probabilities per option (+ confidence)
App->>App: Field threshold check, then act or escalate
All questions about one state travel in one request; whether they share a forward pass depends on the implementation. The application, not the model, decides what each probability means for control flow.
Why it is fast
Generation latency grows with output length. A decision model does not decode an answer, so latency is dominated by reading the input. Grouping questions about one state in a request saves network round trips, and implementations that read the state once also avoid repeated prefill. Published medians range from tens of milliseconds for small models to around half a second for the largest hosted ones, compared with seconds for reasoning LLM calls. Treat specific numbers as vendor-reported and measure on your own traffic.
Architecture
A decision model is rarely a whole system. It is a component on the hot path of a larger pipeline, usually in front of an LLM and a human queue.
Diagram: Decision model inside an AI application
flowchart TB
IN[Inbound event or agent step] --> DM["Decision model:<br/>typed questions"]
DM --> TH{{"Decision signal meets<br/>calibrated threshold?"}}
TH -->|Yes| POL{{"Policy allows<br/>this action?"}}
TH -->|No| LLM[LLM with structured outputs]
LLM --> TH2{{"Valid answer and<br/>threshold met?"}}
TH2 -->|Yes| POL
TH2 -->|No| HQ[Human review queue]
POL -->|Yes| ACT["Deterministic action:<br/>route, tag, approve, skip"]
POL -->|No| HQ
ACT --> LOG[("Decision log: inputs,<br/>probabilities, path,<br/>outcome")]
HQ --> LOG
LOG --> EV["Offline evaluation<br/>and threshold tuning"]
EV -.-> TH
The decision model is one component: thresholds decide whether its answer is trusted, a policy check outside every model decides whether the action is allowed, and logged outcomes feed back into thresholds.
| Component | Responsibility |
|---|---|
| Question schema | Versioned definitions of questions, options, and wording. Treat it like an API contract. |
| Decision model | Scores options and returns distributions. Hosted (TypeSafe, Workers AI) or self-hosted (llama.cpp, SGLang, Ollama). |
| Threshold policy | Per-field cut-offs on a named signal (top probability, P(true)), tuned on labeled data and stricter for risky actions. |
| Policy check | Deterministic rules and permissions that apply whichever model answered. Probability never substitutes for them. |
| Fallback model | An LLM that takes uncertain or out-of-scope cases, often with the same schema via structured outputs. |
| Human queue | The final path for below-threshold, disallowed, or high-risk decisions. See human-in-the-loop. |
| Decision log | Stores state hash, schema version, probabilities, chosen path, and eventual ground truth. |
| Evaluation loop | Measures accuracy and calibration per question, and re-tunes thresholds when data or models change. |
The same layout works inside an agent. Replace "inbound event" with "proposed next step" and the decision model becomes a gate: does this tool call stay within the user's request?, is the task complete?, should we retry? The LLM still plans and writes; the decision model checks.
Step-by-Step Flow
- Identify bounded decisions. List every place your system picks from a known set: routes, categories, severity levels, yes/no checks. Anything that must produce new text stays with an LLM.
- Write the questions. For each decision, write clear instructions and describe each option. Option descriptions matter: "billing — payments, charges, refunds, invoices" outperforms a bare "billing".
- Group questions by state. Questions about the same input go in one request, which saves round trips and, on implementations that read the state once, repeated prefill.
- Collect a labeled evaluation set. A few hundred real examples per question, labeled by people who own the decision, is a reasonable start.
- Measure accuracy and calibration. Run the set through one or more decision models and through your current LLM path. Record accuracy, expected calibration error or Brier score, and latency.
- Set thresholds per question. Choose the signal you will threshold (top probability, P(true), or a provider's
confidence) and the cut-off above which the decision model's answer is accurate enough for that action. Irreversible actions need higher thresholds. - Wire the fallback. Below threshold, send the case to an LLM or a person. Keep the schema identical so downstream code does not care which path answered.
- Log every decision. Store probabilities, path taken, schema version, and model version. Join with ground truth when it becomes known.
- Monitor drift. Track the share of traffic above threshold, accuracy on sampled reviews, and calibration over time. A falling automation rate is often the first sign of drift.
- Re-tune on change. New options, reworded questions, or a model upgrade invalidate old thresholds. Re-run the evaluation set before deploying.
Real Production Example
The example below triages support messages with a decision model over the /v1/systemone contract. It works against TypeSafe's hosted API, a local llama.cpp or Ollama server, or another compatible endpoint by changing the base URL. Uncertain answers fall back to a caller-supplied LLM path.
import os
from dataclasses import dataclass
from typing import Any, Callable
import httpx
# e.g. https://api.typesafe.ai/v1 or http://localhost:8080/v1
BASE_URL = os.environ["SYSTEM_ONE_BASE_URL"]
API_KEY = os.environ.get("SYSTEM_ONE_API_KEY") # local servers usually need none
MODEL = os.environ.get("SYSTEM_ONE_MODEL", "jev-latest")
SCHEMA_VERSION = "support-triage-v3"
QUESTIONS: dict[str, dict[str, Any]] = {
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Payments, charges, refunds, invoices",
"technical": "Outages, errors, bugs, configuration",
"account": "Login, access, account settings",
"sales": "Plans, upgrades, pricing questions",
},
},
"urgent": {
"type": "noul",
"instructions": "Does this request need a response today?",
"criteria": {
"true": "The customer is blocked, losing money, or has a deadline today.",
"false": "It can wait its turn in the queue.",
},
},
"severity": {
"type": "score",
"instructions": "How severe is the customer impact?",
"criteria": ["No impact", "Minor", "Major", "Critical"],
},
}
# Tuned on a labeled set; stricter where a wrong answer is costly. Each
# cut-off is compared with a model probability, not with a server
# `confidence` field, whose formula varies between implementations.
MIN_TOP_PROBABILITY = {"team": 0.80, "severity": 0.60}
# noul returns P(true). Act on True at or above the upper cut-off, on False
# at or below the lower one, and escalate anything in between.
NOUL_CUTOFFS = {"urgent": (0.10, 0.85)}
@dataclass
class Triage:
team: str
urgent: bool
severity: int # index into QUESTIONS["severity"]["criteria"]
path: str # "decision_model" or "fallback"
signals: dict[str, float] # probability each threshold was checked on
def _choice(answer: dict[str, Any]) -> tuple[str, float]:
"""Predicted label and the probability the model gave it."""
label = answer["choice"]
return label, float(answer["probabilities"][label])
def _score_level(answer: dict[str, Any]) -> tuple[int, float]:
"""Most likely rubric level and its probability.
answer["score"] is the probability-weighted mean level, so 1.5 can mean
a split between two adjacent levels or a spread across all of them.
"""
probabilities = answer["probabilities"]
level = max(probabilities, key=probabilities.get)
return int(level), float(probabilities[level])
def _noul_decision(p_true: float, cutoffs: tuple[float, float]) -> bool | None:
"""True or False when P(true) clears a cut-off, else None (escalate)."""
low, high = cutoffs
if p_true >= high:
return True
if p_true <= low:
return False
return None
def decide(state: str, client: httpx.Client) -> dict[str, Any]:
headers = {"Content-Type": "application/json"}
if API_KEY:
headers["Authorization"] = f"Bearer {API_KEY}"
resp = client.post(
f"{BASE_URL.rstrip('/')}/systemone",
json={"model": MODEL, "state": state, "questions": QUESTIONS},
headers=headers,
timeout=httpx.Timeout(2.0, connect=1.0),
)
resp.raise_for_status()
return resp.json()["answers"]
def triage(
state: str,
client: httpx.Client,
fallback: Callable[[str], Triage],
) -> Triage:
try:
answers = decide(state, client)
team, p_team = _choice(answers["team"])
severity, p_severity = _score_level(answers["severity"])
p_urgent = float(answers["urgent"]["noul"])
except (httpx.HTTPError, KeyError, ValueError):
return fallback(state)
urgent = _noul_decision(p_urgent, NOUL_CUTOFFS["urgent"])
trusted = (
p_team >= MIN_TOP_PROBABILITY["team"]
and p_severity >= MIN_TOP_PROBABILITY["severity"]
and urgent is not None
)
if not trusted:
return fallback(state)
return Triage(
team=team,
urgent=urgent,
severity=severity,
path="decision_model",
signals={"team": p_team, "severity": p_severity, "urgent": p_urgent},
)
Production notes:
- Short timeouts. A decision model on the hot path should answer in well under a second. A slow or failed call falls back rather than blocking.
- Four separate steps. The code selects a prediction (the chosen label, the most likely level, or the side of the
noulcut-offs), measures its reliability with a named probability, decides whether to automate by comparing that probability with a tuned cut-off, and records what the predicted value means. Keeping these apart stops one number from being read as another. - What each threshold does. For
team, the cut-off applies to the probability the model gave the chosen team. Forurgent,noulis P(the request needs a response today), not a confidence score; the system acts on "urgent" only at 0.85 or above and on "not urgent" only at 0.10 or below. Thresholdingmax(p, 1 − p)instead would treat 0.12 and 0.88 as equally trustworthy, which hides the fact that a missed urgent ticket usually costs more than a false alarm. Forseverity, the cut-off applies to the most likely level, not to thescoremean; use the mean for sorting and dashboards, and the level for actions. - All-or-nothing escalation. This example escalates the whole message if any field misses its threshold. Per-field escalation is also valid when fields are independent.
- Same schema on both paths. The fallback returns the same
Triagetype, so downstream routing is identical whichever path answered. Recordpath,signals, andSCHEMA_VERSIONwith every decision. - A server
confidenceis a different signal. You can threshold a provider'sconfidencefield instead, but tune the cut-off on that field from that provider. TypeSafe and Kev rescale the top probability against a uniform answer, Ollama and Laya use normalized entropy forchoice, and most return nothing fornoul. Thresholds do not transfer between backends or model versions.
With Pydantic AI, the same decision can be expressed as an output type. Each field becomes a question, and swapping the model name runs the same agent on a language model for comparison:
from enum import Enum
from typing import Annotated
from pydantic import BaseModel, ConfigDict
from pydantic_ai import Agent, BoolCriteria, UseEnumMemberDocstrings
class Team(UseEnumMemberDocstrings, str, Enum):
billing = "billing"
"""Payments, charges, refunds, invoices."""
technical = "technical"
"""Outages, errors, bugs, configuration."""
class Ticket(BaseModel):
"""Triage a support ticket."""
model_config = ConfigDict(use_attribute_docstrings=True)
team: Team
"""Which team should handle this request?"""
urgent: Annotated[
bool,
BoolCriteria(
true="The customer is blocked or losing money today.",
false="It can wait its turn in the queue.",
),
]
"""Does this request need a response today?"""
agent = Agent("typesafe:jev-latest", output_type=Ticket)
result = agent.run_sync("You charged me twice and my account is overdrawn.")
print(result.output)
print(result.response.provider_details["confidence"])
Pydantic AI documents provider_details["confidence"] as a margin, not a probability that the answer is right; for a boolean field it measures distance from the decision_boolean_threshold setting. Tune any threshold on it the same way. Its SystemOneModel targets any /v1/systemone server (for example system-one:clef against a local Ollama), and its docs describe escalating uncertain steps to a language model behind a FallbackModel.
Design Decisions
| Decision | Option A | Option B | When to choose |
|---|---|---|---|
| Hosted vs self-hosted | Hosted API (TypeSafe, Workers AI) | Self-hosted (llama.cpp, SGLang, Ollama) | Hosted for fastest start; self-host for data residency, cost at volume, or fine-tuning control. |
| Model size | Small encoder or ~1–4B decoder | 9B–27B decoder backbone | Small for narrow, high-throughput checks; larger for nuanced policy questions and long states. |
| Escalation granularity | Whole request | Per question | Whole request when fields interact (team and severity); per question when independent. |
| Threshold strategy | One global cut-off | Per question, per action risk | Per question almost always; a single number hides that some decisions are much harder than others. |
| Fallback target | LLM with structured outputs | Human queue | LLM for volume; human for irreversible, regulated, or high-value actions. |
| Option design | Few broad options | Many fine-grained options | Fewer, well-described options calibrate better. Check provider caps (Ollama allows 2–26 options). |
| Customization | Prompt-level (better descriptions) | Fine-tuning (Kev training scripts, Cloudflare's RL program for design partners) | Fix wording first; fine-tune when you have thousands of labeled decisions in a stable domain. |
Common patterns
- Gate, then generate. A decision model decides whether and where to act; an LLM writes only when needed. Example: classify intent, and only call the reply-drafting LLM for intents that need a reply.
- Confidence cascade. Small decision model → larger decision model or LLM → human, stopping at the first answer that clears its threshold. This is model routing driven by a probability signal you have calibrated.
- Many questions, one state. Ask routing, urgency, sentiment, and policy questions about the same message in one request rather than one LLM call per field.
- Agent step checks. Before an agent executes a tool call, ask
noulquestions such as "Does this action stay within what the user asked for?" and "Is this action reversible?" Block or escalate answers below threshold or on risky actions. This complements guardrails; it does not replace them. - Shadow mode first. Run the decision model alongside the existing LLM path, log both, and compare before letting it act.
Comparisons
Decision model vs LLM with structured outputs
| Dimension | Decision model | LLM + structured outputs |
|---|---|---|
| Mechanism | Scores every listed option | Generates tokens constrained to a schema |
| Output space | Bounded by the supplied labels or rubric | Any schema-valid value, including free strings |
| Uncertainty | Probability per option; calibration varies by model | Not native; self-reported confidence is unreliable |
| Latency | Input-bound; no decode loop | Grows with reasoning and output length |
| Cost | Mostly input; no generated answer tokens | Input + output (+ reasoning) tokens |
| Expressiveness | Choice, score, yes/no only | Extraction, free text, nested objects, explanations |
| Reasoning | No decode steps; weak on arithmetic, planning, chains | Can reason step by step before answering |
| Best for | Routing, triage, gating, policy checks at volume | Extraction, generation, decisions that need multi-step reasoning |
Decision model vs trained classifier
| Dimension | Decision model | Task-specific classifier |
|---|---|---|
| New labels | Add them to the request | Collect data and retrain |
| Instructions | Natural language per question | Implicit in training data |
| Probabilities | Per option, for labels supplied at request time | Fast class probabilities or scores; calibration may need extra work |
| Accuracy in-domain | Good zero-shot; improves with tuning | Often best when data is plentiful and stable |
| Size and cost | Hundreds of millions to tens of billions of parameters | Can be tiny |
| Maintenance | One model, many questions | One model per decision |
Decision model vs zero-shot NLI or embedding classifiers
Zero-shot NLI and embedding-similarity classifiers also accept labels at request time. Decision models differ in three ways: they read natural-language instructions and option descriptions jointly, they support ordinal rubrics and several question types in one request, and they are post-trained specifically for decision accuracy and calibration. A published 300-example pilot (jev-benchmarks) compared Jev 1.13 with the open GLiNER2.5 classifier: Jev was well ahead on AG News (0.91 vs 0.70 accuracy) and Banking77 (0.87 vs 0.61), but on an emotion dataset the accuracy gap was unresolved and Jev was markedly worse calibrated. Its README describes it as "a 300-example pilot, not a leaderboard". Treat results like this as directional, especially since public datasets may appear in training data.
Decision tree: when to use a decision model
Decision tree: Choosing a decision model, an LLM, or a classifier
flowchart TD
A["Does the output need<br/>new text, a date, or<br/>a free-form value?"]
B["Is the answer a known<br/>option, a rubric level,<br/>or yes/no?"]
C["Does it need multi-step<br/>reasoning or arithmetic?"]
D["Do labels change often<br/>or need instructions?"]
E["Do latency, cost, or<br/>per-option probabilities<br/>matter?"]
L1(["LLM + structured outputs"])
L2(["LLM + structured outputs"])
L3(["LLM + structured outputs"])
L4(["LLM + structured outputs"])
CLF(["Trained classifier"])
DM(["Decision model with<br/>LLM or human fallback"])
A -->|Yes| L1
A -->|No| B
B -->|No| L2
B -->|Yes| C
C -->|Yes| L3
C -->|No| D
D -->|"No: stable labels,<br/>plenty of data"| CLF
D -->|Yes| E
E -->|No| L4
E -->|Yes| DM
Reach for a decision model when the answer is bounded, the labels are flexible, and speed or per-option probabilities matter.
Common Mistakes
- Putting the question in the state. The state is the material being judged; the question belongs in the question's instructions. A question embedded in the state is just more text to judge, and the model will not treat it as the instruction.
- Trusting vendor calibration on your data. Calibration curves come from the vendor's distribution. Measure on your own labeled set before choosing thresholds.
- Using one global threshold. Questions differ in difficulty and risk. Tune each one.
- Bare option labels. "Other", "misc", or unexplained abbreviations confuse the model. Describe every option, and make options mutually exclusive.
- Asking for values it cannot express. Dates, amounts, names, and explanations need an LLM. Forcing them into choice lists produces brittle schemas.
- Assuming APIs are interchangeable.
/v1/systemoneimplementations share the core request shape but differ in limits (option counts, question counts, context length, image support),confidenceformulas, usage accounting, and streaming. They also handle long states differently: Workers AI truncates, while Kev, Ollama, and llama.cpp reject oversized input. OpenAI's Decisions API (which names the yes/no typepredicate) and SGLang's native/v1/decisionsare separate interfaces. Test each backend. - Skipping the fallback. A decision model without an escalation path either blocks on uncertainty or silently acts on answers below threshold.
- Treating benchmark leaderboards as proof. Many published comparisons are vendor-run, use LLM-generated reference labels, or have small samples. Use them to shortlist models, not to set thresholds.
Where It Breaks Down
- Reasoning-heavy decisions. Without decode steps there is no room to work through a problem, so decision models are weak at arithmetic, date logic, multi-hop policy reasoning, and planning. Kev's maintainer tracks date arithmetic degrading after decision training, and TypeSafe's own jaggedness notes for Jev 1.13 list numbers, dates, indirect references, overly literal reading, and sensitivity to option order.
- Long states. Context limits vary from about 512 tokens (Laya's base checkpoint) to 32K–64K tokens for larger decoder models, and accuracy can drop well before the hard limit. Check each model's validated context, not just its maximum.
- Out-of-distribution domains. Specialist checkpoints fine-tuned on narrow workflows can underperform the general model elsewhere; Laya's typed-decisions model card says so explicitly.
- Overlapping options. When two options are both plausible, probability splits between them. That is honest, but it pushes more traffic to the fallback. Redesign the options before blaming the model.
- Evaluation circularity. TypeSafe's own workflow evaluations score agreement with labels produced by frontier LLMs, not human ground truth. A decision model can match an LLM's mistakes and look accurate.
- Ecosystem immaturity. The category is weeks old. Model names, API routes, and response fields are still changing. OpenAI's Decisions API is in public beta with one supported model, and the yes/no question type already has three names across providers.
When NOT to Use Decision Models
Skip decision models when:
- The output must be written. Replies, summaries, explanations, and extracted free-text values need an LLM.
- The decision needs visible reasoning. If auditors or users need the rationale, use an LLM that can explain, or log the evidence separately.
- Volume is low. A few hundred decisions a day rarely justify a second model and a calibration process; an LLM with structured outputs is simpler.
- Labels are fixed and data is plentiful. A small trained classifier may be cheaper and more accurate.
- You cannot build a labeled evaluation set. Without one you cannot set thresholds, and an uncalibrated decision model is just a faster guess.
- The action is catastrophic if wrong and there is no human path. Probabilities lower risk; they do not remove it.
Warning
A high probability is not authorization. A model's probability or confidence is evidence for a control decision; it is not permission to perform the action. Keep policy checks, permissions, and human approval for irreversible actions regardless of what the decision model returns.
Running in Production
Best Practice
Start in shadow mode, calibrate per question on your own labels, escalate below threshold with the same schema, and log every decision with its probabilities, path, schema version, and model version.
| Dimension | Guidance |
|---|---|
| Scaling | Batch questions per state. Self-hosted servers (llama.cpp router mode, SGLang, Ollama) can host several decision models and pick per request. |
| Latency | Budget hundreds of milliseconds, not seconds. Set tight client timeouts and fall back on timeout rather than blocking the request. |
| Cost | No generated answer; cost tracks input size and request count (check how each provider reports usage). The saving comes from traffic above threshold. |
| Monitoring | Track automation rate per question, fallback rate, histograms of the thresholded signal, timeouts, and accuracy on sampled human reviews. |
| Evaluation | Accuracy, Brier score or calibration error, and automation rate at a fixed error budget, per question. See Evaluation. |
| Security | The state is untrusted input. Prompt-injection text can sway decisions; apply guardrails and keep authorization outside the model. |
Production checklist
- Questions and options versioned like an API schema
- Labeled evaluation set per question, refreshed with production samples
- Per-question thresholds tied to action risk
- Fallback path (LLM or human) returning the same schema
- Client timeouts and fallback on error
- Decision log with probabilities, path, schema version, and model version
- Shadow-mode comparison before enabling actions
- Drift alerts on automation rate and calibration
- Re-evaluation gate for model upgrades and schema changes
- Human approval retained for irreversible or regulated actions
Related Guides
Concept guides
- Structured Outputs: the generative alternative, and the right tool when values must be written.
- Function Calling: tool selection is a decision; arguments usually still need an LLM.
- Model Routing: decision models give routers a per-option probability to calibrate and threshold in cascades.
- Guardrails: policy enforcement that probabilities alone cannot provide.
- Human-in-the-Loop: where below-threshold decisions should land.
- Evaluation: how to measure accuracy and calibration before trusting thresholds.
- Latency Optimization and Cost Optimization: the main reasons to move decisions off the LLM path.
Rankings
Tools
- llama.cpp · SGLang · Ollama · Pydantic AI
If you understood this topic, read next:
Diagram: Learning path for decision models
flowchart LR
A[LLMs] --> B[Structured outputs]
B --> C[Decision models]
C --> D[Model routing]
C --> E[Guardrails]
D --> F[Evaluation]
Prerequisites: Large Language Models · Structured Outputs · Prompt Engineering
Next topics: Model Routing · Guardrails · Evaluation
Interview Questions
-
What is a decision model and how does it differ from an LLM with structured outputs?
It takes a state and typed questions and returns a bounded decision signal per question, with probability information over the possible outcomes, in one request and without autoregressive decoding. An LLM with structured outputs generates a schema-valid answer token by token and has no native per-option probability. -
What question types does the
/v1/systemoneAPI define?
choice(one label from a supplied list),score(a level on an ordered rubric, returned as a probability-weighted value), andnoul(the probability that a yes/no statement is true). -
Why can't a decision model hallucinate a label?
It has no generation step. It scores only the options you listed, so the chosen label is always one of them. It can still pick the wrong one. -
How would you choose an automation threshold?
Pick the signal you will threshold (top probability, P(true) for yes/no, or a provider'sconfidence). Build a labeled set from real traffic, plot accuracy against that signal per question, and choose the lowest threshold that meets the error budget for that action. For yes/no questions, consider separate cut-offs for acting on true and on false. Stricter for irreversible actions. Re-tune on model or schema changes. -
Where do decision models fail?
Multi-step reasoning, arithmetic and date logic, long states beyond validated context, overlapping options, and domains far from training data. -
How would you use a decision model inside an agent?
As a gate on each step: classify intent, check whether a proposed tool call stays in scope, decide whether the task is complete. The LLM plans and writes; uncertain checks escalate to the LLM or a human. -
Why be careful with published decision-model benchmarks?
Many are vendor-run, some score agreement with LLM-generated labels rather than human ground truth, and sample sizes are often small. Use them to shortlist, then evaluate on your own data. -
When would you still prefer a trained classifier?
When labels are stable, labeled data is plentiful, and you need the smallest, cheapest model for one decision.
Key Takeaways
- Decision models answer typed questions with a bounded decision signal and probability information per question, and no autoregressive decoding; they do not write.
- The distinction from structured outputs is mechanical: scoring a closed set versus generating a schema-valid answer.
choice,score, andnoulcover routing, rubrics, and yes/no checks, with options defined per request.- Probability, a server's
confidence, and a calibrated probability of being correct are different signals; whichever one you threshold must be validated against labeled outcomes on your target workload. - Several models and runtimes adopted TypeSafe's
/v1/systemonerequest shape within weeks, forming a de facto compatibility contract, but limits, response fields, naming, and benchmarks are still settling. - Thresholds must be calibrated per question on your own labeled data, with an LLM or human fallback below them.
- Use decision models to keep bounded decisions off the expensive path, and keep LLMs for reasoning and generation.
FAQs
What is a decision model?
A model that takes a state (text, JSON, or sometimes an image) and typed questions, and returns a bounded decision signal for each question, with probability information over the possible outcomes. It never generates free text.
Is "System One model" the same thing as a decision model?
"System One models" is TypeSafe's name for its models, after Kahneman's fast System 1 thinking, and /v1/systemone is TypeSafe's API. Cloudflare, ggml-org (llama.cpp), and Pydantic AI use "decision models", which is the generic term this guide uses. The underlying idea is the same.
How is this different from asking an LLM for a JSON label?
The LLM generates the label, can in principle produce an unlisted value without strict constraints, and does not give a native probability per option. A decision model scores every option, can only return one you listed, and returns the full distribution.
Can decision models hallucinate?
They cannot invent a label, because there is no generation step. They can still choose the wrong option or be overconfident on unfamiliar data, which is why thresholds and evaluation matter.
Which decision models and servers exist?
As of 2026-10-08: TypeSafe's hosted Jev; OpenAI's Decisions API on gpt-6-luna (public beta); Cloudflare's Clef and Clef-flash (Workers AI, open weights under Apache 2.0); open-source Kev, Laya, and AWS's Strands Decider 2B; Liquid AI's open-weight d1-3B and experimental d1-omni-600M (LFM Open License v1.0, free commercial use under $10M annual revenue); and other models listed by llama.cpp and Ollama. Self-hosting is supported by llama.cpp's /v1/systemone (v0.6.0 and later), SGLang's /v1/decisions and its /v1/systemone-compatible route, and Ollama. Pydantic AI and the Vercel AI SDK provide client APIs, and Vercel AI Gateway routes to several decision models.
Do all implementations use the same API?
Most follow TypeSafe's POST /v1/systemone request shape (state, model, questions with choice/score/noul). Limits, confidence formulas, usage accounting, and long-input handling differ. OpenAI's Decisions API is a separate shape (input, questions with predicate/choice/score), which Vercel AI Gateway also serves, and SGLang exposes its own /v1/decisions. Verify each provider's reference.
How accurate are decision models compared with frontier LLMs?
Vendor reports put the best decision models near mid-tier frontier LLMs on bounded workflow tasks at a small fraction of the cost and latency, and behind the strongest reasoning models. Several of those evaluations use LLM-generated reference labels. Measure on your own data.
Can I fine-tune a decision model?
Yes for open models: Kev publishes training scripts, and Cloudflare announced reinforcement-learning fine-tuning for Clef through a design-partner program. Improve question wording and option descriptions first; fine-tune when you have a large, stable labeled set.
Should a decision model replace my guardrails?
No. It can add useful checks (scope, reversibility, policy categories), but authorization, permissions, and human approval for irreversible actions must live outside any model.
How do I start?
Pick one high-volume bounded decision, write the questions, label a few hundred real examples, run a decision model in shadow mode next to your current path, set a threshold, and add a fallback before letting it act.
References
- TypeSafe — Introducing System One Models and Jev
- TypeSafe docs — Confidence
- TypeSafe docs — Jev 1.13 jaggedness
- Cloudflare — Introducing Clef: open-source decision models
- Cloudflare changelog — Clef on Workers AI
- ggml-org — New in llama.cpp: Decision Models
- llama.cpp v0.6.0 release notes
- SGLang v0.5.21 release notes
- Pydantic AI — Decision models
- Pydantic AI — System One API
- Kev — open-source decision models (GitHub)
- Laya typed-decisions model card (Hugging Face)
- Ollama — Clef
- Ollama API — System One
- OpenAI — DevDay 2026 recap
- OpenAI API docs — Decisions
- Strands Agents — Introducing Strands Decider 2B
- Liquid AI — Open d1 decision models
- Liquid AI d1-3B model card (Hugging Face)
- Liquid AI — LFM Open License v1.0
- Vercel AI SDK — Decisions
- Vercel AI Gateway — OpenAI-compatible Decisions API
- AbdelStark — jev-benchmarks: 300-example pilot, Jev vs GLiNER2.5 (GitHub)
- Yin et al. — Benchmarking Zero-shot Text Classification (2019)
- Guo et al. — On Calibration of Modern Neural Networks (2017)
- Zaratiana et al. — GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface (2025)
- Cascaded Language Models for Cost-Effective Human–AI Decision-Making (2025)