LLM Concepts

Decision Models vs. LLMs Guide

Engineering guide to decision models: models that answer typed questions (pick one, score, yes/no) with a bounded decision signal and probability information over the possible outcomes, instead of generating text. Covers how they differ from structured outputs and classifiers, how they work, the /v1/systemone contract, and production patterns.

45 min readIntermediateLast reviewed: 8 October 2026

Quick Summary

A decision model reads a state plus typed questions and returns a bounded decision signal for each question, with probability information over the possible outcomes, in one request and without generating text.

One Analogy

An LLM with structured outputs is a writer filling in a form; a decision model is a scanner that reads the form and puts a probability next to every box.

Engineering Rule

Act on a decision model only when the signal you threshold clears a cut-off calibrated on your own labeled data, and route everything below it to a language model or a human.

TL;DR

  • A decision model answers questions; it does not write. You send a state (text, JSON, sometimes an image) and typed questions. It returns one bounded decision signal per question with probability information over the possible outcomes, without autoregressive token generation.

  • Three question types cover most decisions. choice picks one label from a list you supply, score places the state on an ordered rubric, and noul returns the probability that a yes/no statement is true. Options are defined per request, so no retraining is needed for new labels. The yes/no type has different names across providers: noul (TypeSafe and compatible servers), predicate (OpenAI), and boolean (Vercel AI SDK).

  • The difference from structured outputs is mechanical, not cosmetic. An LLM with structured outputs generates a JSON answer token by token. A decision model scores every option you listed, usually with a prefill-only pass and no decode loop, so it cannot return a label you did not offer and it gives you a full distribution to threshold.

  • Probability, confidence, and correctness are three different numbers. The distribution is the model's probability for each option. A confidence field, where one exists, is an implementation-specific summary of how concentrated that distribution is. Neither means "probability this answer is right" until you calibrate it on your own labeled data.

  • A de facto compatibility contract has formed around TypeSafe's API. TypeSafe introduced POST /v1/systemone with Jev in September 2026. Within three weeks Cloudflare (Clef), open projects (Kev, Laya), llama.cpp, Ollama, and Pydantic AI supported that request shape, and SGLang shipped a compatible route next to its own /v1/decisions. OpenAI's separate Decisions API (POST /v1/decisions) entered public beta on 2026-10-06, and Vercel's AI Gateway serves both shapes. Limits, response fields, and usage accounting still differ.

  • Use it for routing, triage, gating, and checks, not for anything that must be written. Pair it with a language model: the decision model handles bounded decisions that clear a calibrated threshold, and the LLM or a human takes the uncertain and open-ended cases.

On this page

Why This Matters

Most calls an AI system makes are not requests for prose. They are small decisions: which team owns this ticket, is this message urgent, does this tool call stay inside the user's request, is this document a contract, should the agent retry or stop. In a typical agent loop, those decisions outnumber the steps that actually need generated text.

Teams usually make these decisions with a general-purpose LLM and a schema. That works, but it carries costs that compound at volume:

  • Latency on the hot path. An LLM call that reasons and then emits JSON takes seconds. A router or gate that runs on every request makes the whole product slower.
  • Spend on throwaway tokens. Reasoning tokens and JSON syntax are billed even though the application keeps a single label.
  • No usable uncertainty signal. A generated "urgent": true does not say how sure the model was. Engineers bolt on self-reported confidence scores, which are often poorly calibrated, or run the same prompt several times and vote.
  • Shape failures. Even with constrained decoding, providers differ in support, and prompt-only JSON still fails sometimes.

Decision models target exactly this slice of work. The call returns a distribution over the options you defined, so code can branch on it directly and send the uncertain tail to a stronger model or a person. When the decision model is right often enough above your threshold, much of the traffic never reaches the expensive path.

This matters even if you never deploy one. The pattern behind it, separating bounded decisions from open-ended generation and gating actions on calibrated probability, is a sound way to structure any agent or workflow.

The Problem Decision Models Solve

Consider a support pipeline that must route each inbound message, flag urgency, and score customer frustration before an agent drafts a reply.

With a general LLM, you write a prompt, attach a JSON Schema, and parse the result. You get one label per field. If you want to know whether the model was unsure, you either ask it to rate itself (unreliable) or call it several times (slow and expensive). If the provider's strict mode does not cover a schema feature you use, you add retries.

With a trained classifier, you get fast class probabilities or scores (calibration may take extra work, such as temperature scaling), but only for the labels it was trained on. Adding a "security" team means collecting labeled data and retraining. Each new decision (urgency, frustration, refund intent) is a separate model to build and maintain.

With zero-shot classifiers (NLI models or embedding similarity), you can supply labels at request time, but accuracy drops on nuanced instructions, ordinal rubrics are awkward, and you still build one pipeline per question.

A decision model sits between these:

Need General LLM + schema Trained classifier Decision model
Labels defined at request time Yes No Yes
Natural-language instructions per field Yes No Yes
Probability per option No (not natively) Class probabilities or scores Yes
Can return a label you did not offer Possible without strict No No
Answer generated token by token Yes No No
Several questions about one input One call, one JSON One model each One request, all of them
Open-ended text (summary, reply, reason) Yes No No — use an LLM for this

Neither column's probabilities are guaranteed to be calibrated on your traffic; that is something you measure, whichever model produces them.

The problem it solves is narrow but common: fast, cheap answers with a probability per option, for bounded questions whose options change faster than you can retrain a classifier.

How We Got Here

Diagram: Evolution of typed decisions with models

flowchart TB
    A["Pre-2019<br/>Classifier per label set"]
    B["2019<br/>Zero-shot classification<br/>via NLI entailment"]
    C["2023–2024<br/>LLM JSON mode and<br/>schema-constrained output"]
    D["2025 research<br/>Labels set at request time;<br/>confidence-based deferral"]
    E["Sep 15, 2026<br/>TypeSafe ships Jev and<br/>POST /v1/systemone"]
    F["Sep 29 – Oct 2, 2026<br/>OpenAI Decisions preview;<br/>Kev 1.0, Clef, Strands Decider,<br/>SGLang, Ollama support"]
    G["Oct 5–6, 2026<br/>llama.cpp v0.6.0;<br/>OpenAI Decisions API<br/>public beta"]
    A --> B --> C --> D --> E --> F --> G

The recent development is not classification itself, but a model and API pattern built for request-time typed decisions, per-option probabilities, and non-autoregressive inference.

Classification is one of the oldest jobs in machine learning. Through the 2010s, teams trained a model per label set: spam filters, intent classifiers, sentiment models. They were fast and cheap, though modern neural classifiers often needed recalibration before their scores could be read as probabilities (Guo et al., 2017), and they were rigid.

Around 2019, researchers showed that natural language inference (NLI) models could classify text zero-shot by testing whether "This text is about {label}" is entailed. Embedding similarity offered another route. Both let engineers define labels at request time, at some cost in accuracy.

Large language models then absorbed most of this work. With JSON mode and later schema-constrained decoding, an LLM could classify, extract, and score in one call with rich instructions. The trade-off was latency, cost, and the absence of a native probability per option.

By 2025, two research threads had made the ingredients familiar without defining the category. Schema-driven encoders such as GLiNER2 (EMNLP 2025) classified text against labels supplied at inference time, several tasks per call, on a CPU-sized model. Work on cascades and deferral, such as Cascaded Language Models for Cost-Effective Human–AI Decision-Making, used calibrated confidence to decide whether a small model answers, a larger model takes over, or a person reviews. Convai Innovations, which later published Laya, released confidence-aware routing research the same year. None of this work proposed today's decision-model API.

In September 2026, TypeSafe AI released Jev and called it the first "System One model", after Daniel Kahneman's fast, intuitive System 1. TypeSafe described a model that "gives up string generation" in exchange for speed and typed answers with probabilities, served over POST /v1/systemone.

Note

Terminology. This guide uses decision model as the generic engineering term, as Cloudflare, ggml-org (llama.cpp), and Pydantic AI do. System One is TypeSafe's name for its model class, and /v1/systemone is TypeSafe's API, which other projects implement for compatibility. It is not an industry standard.

Within weeks:

  • Jared Palmer released Kev 1.0, an open-source family from 0.8B to 27B parameters that implements the same API. The smaller models are LoRA adapters plus a pointer head on frozen Qwen 3.5 bases; Kev-27B fine-tunes every weight of Qwen3.8-27B.
  • Convai Innovations published Laya, an encoder-based decision model built on ModernBERT that answers the same three question types.
  • Cloudflare released Clef and Clef-flash on Workers AI with open weights under Apache 2.0. Cloudflare says Clef follows the System One API, so a Jev integration can switch by changing the endpoint and model.
  • AWS's Strands Labs released Strands Decider 2B, an open-source model that replaces the language-model head of Qwen3.5-2B with a small pointer head and fine-tunes the backbone with a LoRA adapter. The weights, training data, and scripts are published.
  • Liquid AI released open-weight d1-3B (2026-10-07), built on LFM2.5-VL-3B, which answers typed questions over text and images with a 32K context in one forward pass and zero output tokens. Liquid reports a Decision Index 0.2.1 score of 48.57, the best among models under 10B parameters, and latency of 16 ms on Jetson AGX Thor and 50 ms on Jetson Orin Nano. An experimental d1-omni-600M, built on LFM2.5-Encoder-350M, accepts text with an image or text with audio. llama.cpp supported both on day one. The weights use the LFM Open License v1.0, which allows free commercial use only for organizations under $10M in annual revenue, so check the terms before production use.
  • llama.cpp shipped a /v1/systemone server endpoint in its v0.6.0 release (2026-10-05), with Clef (text and vision) and Nimble among the supported models. SGLang v0.5.21 shipped /v1/decisions, which turns an existing LLM or VLM into a classifier and scorer, plus a System One–compatible route. Ollama added decision models over /v1/systemone, and Pydantic AI added DecisionModel and SystemOneModel classes.
  • Vercel added an experimental experimental_decide function to the AI SDK (ai 7.0.128 or later) and decision models to AI Gateway, which also exposes an OpenAI-compatible /decisions endpoint.
  • OpenAI previewed a Decisions API at DevDay and released it in public beta on 2026-10-06. It uses a dedicated POST /v1/decisions endpoint, supports only gpt-6-luna for now, and bills input tokens only. OpenAI says it expects general availability in the coming weeks.

That pace is the reason for caution as much as interest. The engineering idea is solid and multiply implemented. The benchmarks, naming, and long-term API shape are still settling.

What Is a Decision Model?

A decision model is a model that takes a state and a set of typed questions, and returns a bounded decision signal for each question, with probability information over the possible outcomes. It does not generate free text. A choice answer selects one option from the supplied set, a noul answer is P(true) for the supplied yes/no statement, and a score answer is a probability-weighted value derived from the supplied ordered levels, so it can fall between them.

Inputs

Input What it is
State The material to judge: a string, or a JSON object or array (chat logs, records, app state). Some models also accept images.
Questions A map of named questions. You choose the keys; answers come back under the same keys.
Instructions Natural-language wording for each question ("Is this support request urgent?").
Criteria The options: labels with optional descriptions, ordered rubric levels, or true/false wording.

The three question types

TypeSafe's /v1/systemone API defines three primitives, and the compatible implementations (Laya, Ollama, llama.cpp, Clef, Kev) use the same three:

Type Asks Criteria shape Main answer fields
choice Which one of these options? Map of label → description (or null), or a label list choice (argmax label), probabilities, usually confidence*
score Where on this ordered rubric? Ordered list of level descriptions, lowest first score (probability-weighted mean level), probabilities, legend, confidence*
noul Is this statement true? Optional { "true": "...", "false": "..." } wording noul — the model's probability that the statement is true

* confidence is a summary of the distribution whose formula differs by implementation; it is not the probability that the answer is correct. Most implementations return no confidence for noul.

OpenAI's Decisions API uses the same three ideas with different names and fields: predicate returns probability (the estimate that the condition is true), while choice and score return probabilities and confidence. Images must be sent as inline base64 data URLs. The Vercel AI SDK calls the yes/no type boolean. Map field names explicitly when you switch providers.

A choice question generalizes multi-class classification with labels supplied at request time. A score question handles ordinal judgments (severity, frustration, quality) where "Major" is closer to "Critical" than to "No impact". A noul question returns a single probability for a yes/no statement; how well that probability is calibrated depends on the model and on your data.

Outputs

The response contains an answer per question and a usage block. There is no autoregressive decode, so nothing in the answer is generated text. Usage accounting is provider-specific: llama.cpp and Laya report output_tokens: 0, while TypeSafe's API reference shows small non-zero counts and Ollama documents the field as tokens produced internally for scoring. Read each provider's reference before parsing fields or estimating cost.

Three numbers in a response are easy to conflate:

  • Probability is the distribution the model assigns over the options: probabilities for choice and score, and noul as P(true). It describes what the model believes, not whether the model is right.
  • Confidence is an implementation-specific summary of how concentrated that distribution is. TypeSafe and Kev rescale the top probability (or the spread around the mean level) against a uniform answer; Ollama and Laya use normalized entropy for choice; Pydantic AI reports a margin from its decision threshold. Values from different backends are not comparable, and none of them is P(correct).
  • Calibrated correctness probability is a signal you validated against labeled outcomes on your target traffic, for example "of answers with top probability at or above 0.9, 97% were right". No API returns this by default; you build it by measuring one of the first two signals against your own labels.

You can threshold top probability, P(true), a provider's confidence, or another explicitly defined signal, but none of them is a calibrated probability by default. Any signal used for automation must be validated against labeled outcomes on your target workload.

What a decision model is not

  • Not a chat model. It cannot write a reply, summary, or explanation.
  • Not an extractor of arbitrary values. It cannot return a free-form string, date, or unbounded number. Those still need an LLM with structured outputs.
  • Not a policy engine. It estimates probabilities. Your code decides what they mean and what to do with them.

How Decision Models Work

Scoring instead of generating

An autoregressive LLM produces an answer one token at a time, each step conditioned on the previous ones. Even with constrained decoding, it is sampling a path through the space of valid outputs.

A decision model instead treats each question's options as a closed set and scores all of them. Implementations published so far follow two broad designs.

Decoder backbone, prefill only, plus a scoring head. A causal language model reads the state, question, and options in prefill only, with no decode loop. A small head then scores each option against the question:

  • Kev's README describes a Qwen backbone run prefill-only, LoRA adapters on the smaller checkpoints (Kev-27B is fully fine-tuned), and a pointer head that scores each option's closing token against the question's final token. A softmax gives the distribution.
  • Cloudflare describes Clef as a frozen Qwen backbone (Qwen3.8-27B for Clef, Qwen3.5-9B for Clef-flash) with a routing head and low-rank adapters. Each valid option gathers evidence from the prompt, fields cross-attend to one another, and the decision step is non-autoregressive.

Encoder backbone with option markers. Laya places a [MASK] marker per option, reads the whole sequence bidirectionally with a ModernBERT encoder, refines it with a small decision head, and reads one logit per marker. A learned type embedding tells the model whether it is answering choice, score, or noul. This design is far smaller (hundreds of millions of parameters) with a shorter context window.

In both designs, the model cannot produce an option you did not list, because there is no generation step to produce one.

Several questions travel in one request, but how they are executed varies. TypeSafe evaluates questions "in parallel and in isolation"; llama.cpp answers them independently, with some models reading the state only once; Laya batches all questions into one forward pass; current Kev models run each question as its own sequence. Clef is the notable exception: Cloudflare says its fields cross-attend, so answers can inform one another. Unless your provider documents otherwise, assume questions are independent and one answer cannot see another.

Calibration

A probability is only useful for automation if it means what it says: of all answers given at 0.9, about 90% should be right. Several decision-model builders report treating calibration as a training target:

  • Proper scoring rules. Clef's post-training pairs label-smoothed cross-entropy with a Brier loss. Laya's specialist checkpoint uses a reward built from strictly proper scoring rules, which are maximized only by honest probabilities.
  • Temperature scaling. Kev fits a temperature per checkpoint on held-out data and ships it as the default. Its README notes that a single temperature cannot change which answers are ranked as more or less certain.
  • Reinforcement Learning for Calibrated Decisions (RLCD). TypeSafe introduced the term for its training method; Cloudflare and Laya describe their own RLCD variants that reward calibrated distributions and give partial credit to adjacent ordinal levels.

The results are uneven. Laya's model card reports an expected calibration error of 0.213 for its base checkpoint and advises treating its confidence as uncalibrated, and a published 300-example pilot found Jev worse calibrated than a GLiNER2.5 baseline on one emotion dataset. Calibration measured on a vendor's data does not transfer automatically to yours. Re-measure on your own labeled set before setting thresholds.

Diagram: One decision-model request

sequenceDiagram
    participant App
    participant DM as Decision model
    participant Head as Scoring head
    App->>DM: state + typed questions
    DM->>DM: Prefill only, no decoding
    DM->>Head: Option representations
    Head-->>DM: Logit per option, softmax
    DM-->>App: Probabilities per option (+ confidence)
    App->>App: Field threshold check, then act or escalate

All questions about one state travel in one request; whether they share a forward pass depends on the implementation. The application, not the model, decides what each probability means for control flow.

Why it is fast

Generation latency grows with output length. A decision model does not decode an answer, so latency is dominated by reading the input. Grouping questions about one state in a request saves network round trips, and implementations that read the state once also avoid repeated prefill. Published medians range from tens of milliseconds for small models to around half a second for the largest hosted ones, compared with seconds for reasoning LLM calls. Treat specific numbers as vendor-reported and measure on your own traffic.

Architecture

A decision model is rarely a whole system. It is a component on the hot path of a larger pipeline, usually in front of an LLM and a human queue.

Diagram: Decision model inside an AI application

flowchart TB
    IN[Inbound event or agent step] --> DM["Decision model:<br/>typed questions"]
    DM --> TH{{"Decision signal meets<br/>calibrated threshold?"}}
    TH -->|Yes| POL{{"Policy allows<br/>this action?"}}
    TH -->|No| LLM[LLM with structured outputs]
    LLM --> TH2{{"Valid answer and<br/>threshold met?"}}
    TH2 -->|Yes| POL
    TH2 -->|No| HQ[Human review queue]
    POL -->|Yes| ACT["Deterministic action:<br/>route, tag, approve, skip"]
    POL -->|No| HQ
    ACT --> LOG[("Decision log: inputs,<br/>probabilities, path,<br/>outcome")]
    HQ --> LOG
    LOG --> EV["Offline evaluation<br/>and threshold tuning"]
    EV -.-> TH

The decision model is one component: thresholds decide whether its answer is trusted, a policy check outside every model decides whether the action is allowed, and logged outcomes feed back into thresholds.

Component Responsibility
Question schema Versioned definitions of questions, options, and wording. Treat it like an API contract.
Decision model Scores options and returns distributions. Hosted (TypeSafe, Workers AI) or self-hosted (llama.cpp, SGLang, Ollama).
Threshold policy Per-field cut-offs on a named signal (top probability, P(true)), tuned on labeled data and stricter for risky actions.
Policy check Deterministic rules and permissions that apply whichever model answered. Probability never substitutes for them.
Fallback model An LLM that takes uncertain or out-of-scope cases, often with the same schema via structured outputs.
Human queue The final path for below-threshold, disallowed, or high-risk decisions. See human-in-the-loop.
Decision log Stores state hash, schema version, probabilities, chosen path, and eventual ground truth.
Evaluation loop Measures accuracy and calibration per question, and re-tunes thresholds when data or models change.

The same layout works inside an agent. Replace "inbound event" with "proposed next step" and the decision model becomes a gate: does this tool call stay within the user's request?, is the task complete?, should we retry? The LLM still plans and writes; the decision model checks.

Step-by-Step Flow

  1. Identify bounded decisions. List every place your system picks from a known set: routes, categories, severity levels, yes/no checks. Anything that must produce new text stays with an LLM.
  2. Write the questions. For each decision, write clear instructions and describe each option. Option descriptions matter: "billing — payments, charges, refunds, invoices" outperforms a bare "billing".
  3. Group questions by state. Questions about the same input go in one request, which saves round trips and, on implementations that read the state once, repeated prefill.
  4. Collect a labeled evaluation set. A few hundred real examples per question, labeled by people who own the decision, is a reasonable start.
  5. Measure accuracy and calibration. Run the set through one or more decision models and through your current LLM path. Record accuracy, expected calibration error or Brier score, and latency.
  6. Set thresholds per question. Choose the signal you will threshold (top probability, P(true), or a provider's confidence) and the cut-off above which the decision model's answer is accurate enough for that action. Irreversible actions need higher thresholds.
  7. Wire the fallback. Below threshold, send the case to an LLM or a person. Keep the schema identical so downstream code does not care which path answered.
  8. Log every decision. Store probabilities, path taken, schema version, and model version. Join with ground truth when it becomes known.
  9. Monitor drift. Track the share of traffic above threshold, accuracy on sampled reviews, and calibration over time. A falling automation rate is often the first sign of drift.
  10. Re-tune on change. New options, reworded questions, or a model upgrade invalidate old thresholds. Re-run the evaluation set before deploying.

Real Production Example

The example below triages support messages with a decision model over the /v1/systemone contract. It works against TypeSafe's hosted API, a local llama.cpp or Ollama server, or another compatible endpoint by changing the base URL. Uncertain answers fall back to a caller-supplied LLM path.

import os
from dataclasses import dataclass
from typing import Any, Callable

import httpx

# e.g. https://api.typesafe.ai/v1 or http://localhost:8080/v1
BASE_URL = os.environ["SYSTEM_ONE_BASE_URL"]
API_KEY = os.environ.get("SYSTEM_ONE_API_KEY")  # local servers usually need none
MODEL = os.environ.get("SYSTEM_ONE_MODEL", "jev-latest")
SCHEMA_VERSION = "support-triage-v3"

QUESTIONS: dict[str, dict[str, Any]] = {
    "team": {
        "type": "choice",
        "instructions": "Which team should handle this request?",
        "criteria": {
            "billing": "Payments, charges, refunds, invoices",
            "technical": "Outages, errors, bugs, configuration",
            "account": "Login, access, account settings",
            "sales": "Plans, upgrades, pricing questions",
        },
    },
    "urgent": {
        "type": "noul",
        "instructions": "Does this request need a response today?",
        "criteria": {
            "true": "The customer is blocked, losing money, or has a deadline today.",
            "false": "It can wait its turn in the queue.",
        },
    },
    "severity": {
        "type": "score",
        "instructions": "How severe is the customer impact?",
        "criteria": ["No impact", "Minor", "Major", "Critical"],
    },
}

# Tuned on a labeled set; stricter where a wrong answer is costly. Each
# cut-off is compared with a model probability, not with a server
# `confidence` field, whose formula varies between implementations.
MIN_TOP_PROBABILITY = {"team": 0.80, "severity": 0.60}
# noul returns P(true). Act on True at or above the upper cut-off, on False
# at or below the lower one, and escalate anything in between.
NOUL_CUTOFFS = {"urgent": (0.10, 0.85)}


@dataclass
class Triage:
    team: str
    urgent: bool
    severity: int  # index into QUESTIONS["severity"]["criteria"]
    path: str  # "decision_model" or "fallback"
    signals: dict[str, float]  # probability each threshold was checked on


def _choice(answer: dict[str, Any]) -> tuple[str, float]:
    """Predicted label and the probability the model gave it."""
    label = answer["choice"]
    return label, float(answer["probabilities"][label])


def _score_level(answer: dict[str, Any]) -> tuple[int, float]:
    """Most likely rubric level and its probability.

    answer["score"] is the probability-weighted mean level, so 1.5 can mean
    a split between two adjacent levels or a spread across all of them.
    """
    probabilities = answer["probabilities"]
    level = max(probabilities, key=probabilities.get)
    return int(level), float(probabilities[level])


def _noul_decision(p_true: float, cutoffs: tuple[float, float]) -> bool | None:
    """True or False when P(true) clears a cut-off, else None (escalate)."""
    low, high = cutoffs
    if p_true >= high:
        return True
    if p_true <= low:
        return False
    return None


def decide(state: str, client: httpx.Client) -> dict[str, Any]:
    headers = {"Content-Type": "application/json"}
    if API_KEY:
        headers["Authorization"] = f"Bearer {API_KEY}"
    resp = client.post(
        f"{BASE_URL.rstrip('/')}/systemone",
        json={"model": MODEL, "state": state, "questions": QUESTIONS},
        headers=headers,
        timeout=httpx.Timeout(2.0, connect=1.0),
    )
    resp.raise_for_status()
    return resp.json()["answers"]


def triage(
    state: str,
    client: httpx.Client,
    fallback: Callable[[str], Triage],
) -> Triage:
    try:
        answers = decide(state, client)
        team, p_team = _choice(answers["team"])
        severity, p_severity = _score_level(answers["severity"])
        p_urgent = float(answers["urgent"]["noul"])
    except (httpx.HTTPError, KeyError, ValueError):
        return fallback(state)

    urgent = _noul_decision(p_urgent, NOUL_CUTOFFS["urgent"])
    trusted = (
        p_team >= MIN_TOP_PROBABILITY["team"]
        and p_severity >= MIN_TOP_PROBABILITY["severity"]
        and urgent is not None
    )
    if not trusted:
        return fallback(state)

    return Triage(
        team=team,
        urgent=urgent,
        severity=severity,
        path="decision_model",
        signals={"team": p_team, "severity": p_severity, "urgent": p_urgent},
    )

Production notes:

  • Short timeouts. A decision model on the hot path should answer in well under a second. A slow or failed call falls back rather than blocking.
  • Four separate steps. The code selects a prediction (the chosen label, the most likely level, or the side of the noul cut-offs), measures its reliability with a named probability, decides whether to automate by comparing that probability with a tuned cut-off, and records what the predicted value means. Keeping these apart stops one number from being read as another.
  • What each threshold does. For team, the cut-off applies to the probability the model gave the chosen team. For urgent, noul is P(the request needs a response today), not a confidence score; the system acts on "urgent" only at 0.85 or above and on "not urgent" only at 0.10 or below. Thresholding max(p, 1 − p) instead would treat 0.12 and 0.88 as equally trustworthy, which hides the fact that a missed urgent ticket usually costs more than a false alarm. For severity, the cut-off applies to the most likely level, not to the score mean; use the mean for sorting and dashboards, and the level for actions.
  • All-or-nothing escalation. This example escalates the whole message if any field misses its threshold. Per-field escalation is also valid when fields are independent.
  • Same schema on both paths. The fallback returns the same Triage type, so downstream routing is identical whichever path answered. Record path, signals, and SCHEMA_VERSION with every decision.
  • A server confidence is a different signal. You can threshold a provider's confidence field instead, but tune the cut-off on that field from that provider. TypeSafe and Kev rescale the top probability against a uniform answer, Ollama and Laya use normalized entropy for choice, and most return nothing for noul. Thresholds do not transfer between backends or model versions.

With Pydantic AI, the same decision can be expressed as an output type. Each field becomes a question, and swapping the model name runs the same agent on a language model for comparison:

from enum import Enum
from typing import Annotated

from pydantic import BaseModel, ConfigDict
from pydantic_ai import Agent, BoolCriteria, UseEnumMemberDocstrings


class Team(UseEnumMemberDocstrings, str, Enum):
    billing = "billing"
    """Payments, charges, refunds, invoices."""

    technical = "technical"
    """Outages, errors, bugs, configuration."""


class Ticket(BaseModel):
    """Triage a support ticket."""

    model_config = ConfigDict(use_attribute_docstrings=True)

    team: Team
    """Which team should handle this request?"""

    urgent: Annotated[
        bool,
        BoolCriteria(
            true="The customer is blocked or losing money today.",
            false="It can wait its turn in the queue.",
        ),
    ]
    """Does this request need a response today?"""


agent = Agent("typesafe:jev-latest", output_type=Ticket)
result = agent.run_sync("You charged me twice and my account is overdrawn.")
print(result.output)
print(result.response.provider_details["confidence"])

Pydantic AI documents provider_details["confidence"] as a margin, not a probability that the answer is right; for a boolean field it measures distance from the decision_boolean_threshold setting. Tune any threshold on it the same way. Its SystemOneModel targets any /v1/systemone server (for example system-one:clef against a local Ollama), and its docs describe escalating uncertain steps to a language model behind a FallbackModel.

Design Decisions

Decision Option A Option B When to choose
Hosted vs self-hosted Hosted API (TypeSafe, Workers AI) Self-hosted (llama.cpp, SGLang, Ollama) Hosted for fastest start; self-host for data residency, cost at volume, or fine-tuning control.
Model size Small encoder or ~1–4B decoder 9B–27B decoder backbone Small for narrow, high-throughput checks; larger for nuanced policy questions and long states.
Escalation granularity Whole request Per question Whole request when fields interact (team and severity); per question when independent.
Threshold strategy One global cut-off Per question, per action risk Per question almost always; a single number hides that some decisions are much harder than others.
Fallback target LLM with structured outputs Human queue LLM for volume; human for irreversible, regulated, or high-value actions.
Option design Few broad options Many fine-grained options Fewer, well-described options calibrate better. Check provider caps (Ollama allows 2–26 options).
Customization Prompt-level (better descriptions) Fine-tuning (Kev training scripts, Cloudflare's RL program for design partners) Fix wording first; fine-tune when you have thousands of labeled decisions in a stable domain.

Common patterns

  • Gate, then generate. A decision model decides whether and where to act; an LLM writes only when needed. Example: classify intent, and only call the reply-drafting LLM for intents that need a reply.
  • Confidence cascade. Small decision model → larger decision model or LLM → human, stopping at the first answer that clears its threshold. This is model routing driven by a probability signal you have calibrated.
  • Many questions, one state. Ask routing, urgency, sentiment, and policy questions about the same message in one request rather than one LLM call per field.
  • Agent step checks. Before an agent executes a tool call, ask noul questions such as "Does this action stay within what the user asked for?" and "Is this action reversible?" Block or escalate answers below threshold or on risky actions. This complements guardrails; it does not replace them.
  • Shadow mode first. Run the decision model alongside the existing LLM path, log both, and compare before letting it act.

Comparisons

Decision model vs LLM with structured outputs

Dimension Decision model LLM + structured outputs
Mechanism Scores every listed option Generates tokens constrained to a schema
Output space Bounded by the supplied labels or rubric Any schema-valid value, including free strings
Uncertainty Probability per option; calibration varies by model Not native; self-reported confidence is unreliable
Latency Input-bound; no decode loop Grows with reasoning and output length
Cost Mostly input; no generated answer tokens Input + output (+ reasoning) tokens
Expressiveness Choice, score, yes/no only Extraction, free text, nested objects, explanations
Reasoning No decode steps; weak on arithmetic, planning, chains Can reason step by step before answering
Best for Routing, triage, gating, policy checks at volume Extraction, generation, decisions that need multi-step reasoning

Decision model vs trained classifier

Dimension Decision model Task-specific classifier
New labels Add them to the request Collect data and retrain
Instructions Natural language per question Implicit in training data
Probabilities Per option, for labels supplied at request time Fast class probabilities or scores; calibration may need extra work
Accuracy in-domain Good zero-shot; improves with tuning Often best when data is plentiful and stable
Size and cost Hundreds of millions to tens of billions of parameters Can be tiny
Maintenance One model, many questions One model per decision

Decision model vs zero-shot NLI or embedding classifiers

Zero-shot NLI and embedding-similarity classifiers also accept labels at request time. Decision models differ in three ways: they read natural-language instructions and option descriptions jointly, they support ordinal rubrics and several question types in one request, and they are post-trained specifically for decision accuracy and calibration. A published 300-example pilot (jev-benchmarks) compared Jev 1.13 with the open GLiNER2.5 classifier: Jev was well ahead on AG News (0.91 vs 0.70 accuracy) and Banking77 (0.87 vs 0.61), but on an emotion dataset the accuracy gap was unresolved and Jev was markedly worse calibrated. Its README describes it as "a 300-example pilot, not a leaderboard". Treat results like this as directional, especially since public datasets may appear in training data.

Decision tree: when to use a decision model

Decision tree: Choosing a decision model, an LLM, or a classifier

flowchart TD
    A["Does the output need<br/>new text, a date, or<br/>a free-form value?"]
    B["Is the answer a known<br/>option, a rubric level,<br/>or yes/no?"]
    C["Does it need multi-step<br/>reasoning or arithmetic?"]
    D["Do labels change often<br/>or need instructions?"]
    E["Do latency, cost, or<br/>per-option probabilities<br/>matter?"]
    L1(["LLM + structured outputs"])
    L2(["LLM + structured outputs"])
    L3(["LLM + structured outputs"])
    L4(["LLM + structured outputs"])
    CLF(["Trained classifier"])
    DM(["Decision model with<br/>LLM or human fallback"])
    A -->|Yes| L1
    A -->|No| B
    B -->|No| L2
    B -->|Yes| C
    C -->|Yes| L3
    C -->|No| D
    D -->|"No: stable labels,<br/>plenty of data"| CLF
    D -->|Yes| E
    E -->|No| L4
    E -->|Yes| DM

Reach for a decision model when the answer is bounded, the labels are flexible, and speed or per-option probabilities matter.

Common Mistakes

  1. Putting the question in the state. The state is the material being judged; the question belongs in the question's instructions. A question embedded in the state is just more text to judge, and the model will not treat it as the instruction.
  2. Trusting vendor calibration on your data. Calibration curves come from the vendor's distribution. Measure on your own labeled set before choosing thresholds.
  3. Using one global threshold. Questions differ in difficulty and risk. Tune each one.
  4. Bare option labels. "Other", "misc", or unexplained abbreviations confuse the model. Describe every option, and make options mutually exclusive.
  5. Asking for values it cannot express. Dates, amounts, names, and explanations need an LLM. Forcing them into choice lists produces brittle schemas.
  6. Assuming APIs are interchangeable. /v1/systemone implementations share the core request shape but differ in limits (option counts, question counts, context length, image support), confidence formulas, usage accounting, and streaming. They also handle long states differently: Workers AI truncates, while Kev, Ollama, and llama.cpp reject oversized input. OpenAI's Decisions API (which names the yes/no type predicate) and SGLang's native /v1/decisions are separate interfaces. Test each backend.
  7. Skipping the fallback. A decision model without an escalation path either blocks on uncertainty or silently acts on answers below threshold.
  8. Treating benchmark leaderboards as proof. Many published comparisons are vendor-run, use LLM-generated reference labels, or have small samples. Use them to shortlist models, not to set thresholds.

Where It Breaks Down

  • Reasoning-heavy decisions. Without decode steps there is no room to work through a problem, so decision models are weak at arithmetic, date logic, multi-hop policy reasoning, and planning. Kev's maintainer tracks date arithmetic degrading after decision training, and TypeSafe's own jaggedness notes for Jev 1.13 list numbers, dates, indirect references, overly literal reading, and sensitivity to option order.
  • Long states. Context limits vary from about 512 tokens (Laya's base checkpoint) to 32K–64K tokens for larger decoder models, and accuracy can drop well before the hard limit. Check each model's validated context, not just its maximum.
  • Out-of-distribution domains. Specialist checkpoints fine-tuned on narrow workflows can underperform the general model elsewhere; Laya's typed-decisions model card says so explicitly.
  • Overlapping options. When two options are both plausible, probability splits between them. That is honest, but it pushes more traffic to the fallback. Redesign the options before blaming the model.
  • Evaluation circularity. TypeSafe's own workflow evaluations score agreement with labels produced by frontier LLMs, not human ground truth. A decision model can match an LLM's mistakes and look accurate.
  • Ecosystem immaturity. The category is weeks old. Model names, API routes, and response fields are still changing. OpenAI's Decisions API is in public beta with one supported model, and the yes/no question type already has three names across providers.

When NOT to Use Decision Models

Skip decision models when:

  1. The output must be written. Replies, summaries, explanations, and extracted free-text values need an LLM.
  2. The decision needs visible reasoning. If auditors or users need the rationale, use an LLM that can explain, or log the evidence separately.
  3. Volume is low. A few hundred decisions a day rarely justify a second model and a calibration process; an LLM with structured outputs is simpler.
  4. Labels are fixed and data is plentiful. A small trained classifier may be cheaper and more accurate.
  5. You cannot build a labeled evaluation set. Without one you cannot set thresholds, and an uncalibrated decision model is just a faster guess.
  6. The action is catastrophic if wrong and there is no human path. Probabilities lower risk; they do not remove it.

Warning

A high probability is not authorization. A model's probability or confidence is evidence for a control decision; it is not permission to perform the action. Keep policy checks, permissions, and human approval for irreversible actions regardless of what the decision model returns.

Running in Production

Best Practice

Start in shadow mode, calibrate per question on your own labels, escalate below threshold with the same schema, and log every decision with its probabilities, path, schema version, and model version.

Dimension Guidance
Scaling Batch questions per state. Self-hosted servers (llama.cpp router mode, SGLang, Ollama) can host several decision models and pick per request.
Latency Budget hundreds of milliseconds, not seconds. Set tight client timeouts and fall back on timeout rather than blocking the request.
Cost No generated answer; cost tracks input size and request count (check how each provider reports usage). The saving comes from traffic above threshold.
Monitoring Track automation rate per question, fallback rate, histograms of the thresholded signal, timeouts, and accuracy on sampled human reviews.
Evaluation Accuracy, Brier score or calibration error, and automation rate at a fixed error budget, per question. See Evaluation.
Security The state is untrusted input. Prompt-injection text can sway decisions; apply guardrails and keep authorization outside the model.

Production checklist

  • Questions and options versioned like an API schema
  • Labeled evaluation set per question, refreshed with production samples
  • Per-question thresholds tied to action risk
  • Fallback path (LLM or human) returning the same schema
  • Client timeouts and fallback on error
  • Decision log with probabilities, path, schema version, and model version
  • Shadow-mode comparison before enabling actions
  • Drift alerts on automation rate and calibration
  • Re-evaluation gate for model upgrades and schema changes
  • Human approval retained for irreversible or regulated actions

Concept guides

Rankings

Tools

If you understood this topic, read next:

Diagram: Learning path for decision models

flowchart LR
    A[LLMs] --> B[Structured outputs]
    B --> C[Decision models]
    C --> D[Model routing]
    C --> E[Guardrails]
    D --> F[Evaluation]

Prerequisites: Large Language Models · Structured Outputs · Prompt Engineering

Next topics: Model Routing · Guardrails · Evaluation

Interview Questions

  1. What is a decision model and how does it differ from an LLM with structured outputs?
    It takes a state and typed questions and returns a bounded decision signal per question, with probability information over the possible outcomes, in one request and without autoregressive decoding. An LLM with structured outputs generates a schema-valid answer token by token and has no native per-option probability.

  2. What question types does the /v1/systemone API define?
    choice (one label from a supplied list), score (a level on an ordered rubric, returned as a probability-weighted value), and noul (the probability that a yes/no statement is true).

  3. Why can't a decision model hallucinate a label?
    It has no generation step. It scores only the options you listed, so the chosen label is always one of them. It can still pick the wrong one.

  4. How would you choose an automation threshold?
    Pick the signal you will threshold (top probability, P(true) for yes/no, or a provider's confidence). Build a labeled set from real traffic, plot accuracy against that signal per question, and choose the lowest threshold that meets the error budget for that action. For yes/no questions, consider separate cut-offs for acting on true and on false. Stricter for irreversible actions. Re-tune on model or schema changes.

  5. Where do decision models fail?
    Multi-step reasoning, arithmetic and date logic, long states beyond validated context, overlapping options, and domains far from training data.

  6. How would you use a decision model inside an agent?
    As a gate on each step: classify intent, check whether a proposed tool call stays in scope, decide whether the task is complete. The LLM plans and writes; uncertain checks escalate to the LLM or a human.

  7. Why be careful with published decision-model benchmarks?
    Many are vendor-run, some score agreement with LLM-generated labels rather than human ground truth, and sample sizes are often small. Use them to shortlist, then evaluate on your own data.

  8. When would you still prefer a trained classifier?
    When labels are stable, labeled data is plentiful, and you need the smallest, cheapest model for one decision.

Key Takeaways

  • Decision models answer typed questions with a bounded decision signal and probability information per question, and no autoregressive decoding; they do not write.
  • The distinction from structured outputs is mechanical: scoring a closed set versus generating a schema-valid answer.
  • choice, score, and noul cover routing, rubrics, and yes/no checks, with options defined per request.
  • Probability, a server's confidence, and a calibrated probability of being correct are different signals; whichever one you threshold must be validated against labeled outcomes on your target workload.
  • Several models and runtimes adopted TypeSafe's /v1/systemone request shape within weeks, forming a de facto compatibility contract, but limits, response fields, naming, and benchmarks are still settling.
  • Thresholds must be calibrated per question on your own labeled data, with an LLM or human fallback below them.
  • Use decision models to keep bounded decisions off the expensive path, and keep LLMs for reasoning and generation.

FAQs

What is a decision model?

A model that takes a state (text, JSON, or sometimes an image) and typed questions, and returns a bounded decision signal for each question, with probability information over the possible outcomes. It never generates free text.

Is "System One model" the same thing as a decision model?

"System One models" is TypeSafe's name for its models, after Kahneman's fast System 1 thinking, and /v1/systemone is TypeSafe's API. Cloudflare, ggml-org (llama.cpp), and Pydantic AI use "decision models", which is the generic term this guide uses. The underlying idea is the same.

How is this different from asking an LLM for a JSON label?

The LLM generates the label, can in principle produce an unlisted value without strict constraints, and does not give a native probability per option. A decision model scores every option, can only return one you listed, and returns the full distribution.

Can decision models hallucinate?

They cannot invent a label, because there is no generation step. They can still choose the wrong option or be overconfident on unfamiliar data, which is why thresholds and evaluation matter.

Which decision models and servers exist?

As of 2026-10-08: TypeSafe's hosted Jev; OpenAI's Decisions API on gpt-6-luna (public beta); Cloudflare's Clef and Clef-flash (Workers AI, open weights under Apache 2.0); open-source Kev, Laya, and AWS's Strands Decider 2B; Liquid AI's open-weight d1-3B and experimental d1-omni-600M (LFM Open License v1.0, free commercial use under $10M annual revenue); and other models listed by llama.cpp and Ollama. Self-hosting is supported by llama.cpp's /v1/systemone (v0.6.0 and later), SGLang's /v1/decisions and its /v1/systemone-compatible route, and Ollama. Pydantic AI and the Vercel AI SDK provide client APIs, and Vercel AI Gateway routes to several decision models.

Do all implementations use the same API?

Most follow TypeSafe's POST /v1/systemone request shape (state, model, questions with choice/score/noul). Limits, confidence formulas, usage accounting, and long-input handling differ. OpenAI's Decisions API is a separate shape (input, questions with predicate/choice/score), which Vercel AI Gateway also serves, and SGLang exposes its own /v1/decisions. Verify each provider's reference.

How accurate are decision models compared with frontier LLMs?

Vendor reports put the best decision models near mid-tier frontier LLMs on bounded workflow tasks at a small fraction of the cost and latency, and behind the strongest reasoning models. Several of those evaluations use LLM-generated reference labels. Measure on your own data.

Can I fine-tune a decision model?

Yes for open models: Kev publishes training scripts, and Cloudflare announced reinforcement-learning fine-tuning for Clef through a design-partner program. Improve question wording and option descriptions first; fine-tune when you have a large, stable labeled set.

Should a decision model replace my guardrails?

No. It can add useful checks (scope, reversibility, policy categories), but authorization, permissions, and human approval for irreversible actions must live outside any model.

How do I start?

Pick one high-volume bounded decision, write the questions, label a few hundred real examples, run a decision model in shadow mode next to your current path, set a threshold, and add a fallback before letting it act.

References

Further Reading

Next Topics

Learning Path

Continue Learning

Related Guides

Related companies

  • OpenAI

    Commercial foundation model leader.

Related Tools

ToolCategoryPurposeWebsiteBest For
llama.cpp
Open SourceSelf-hosted
servingHigh-performance C/C++ inference for LLaMA and GGUF models on CPU and GPU.github.comLocal CPU inference
SGLang
Open SourceAPI
servingFast structured generation and serving runtime for LLMs.sgl-project.github.ioStructured generation
Ollama
Open SourceAPI
servingLocal model runner with simple command and HTTP interface.ollama.aiLocal LLM development
PydanticAI
Open SourceAPI
frameworksType-safe Python agent framework with Pydantic validation and structured outputs.ai.pydantic.devType-safe agents

Related Rankings