Model Routing in One Sentence
Model Routing =
- Task and capability match
- Hard policy constraints
- Quality / latency / cost preferences
- Availability
Fallback is not part of that definition. It is a separate, bounded recovery path used when the preferred route cannot be used or fails.
TL;DR
-
Model routing is a decision, not a product. Given a request, the system chooses which eligible model or provider should handle it. A static map such as
classification → small model,summarization → efficient model,complex reasoning → stronger modelis already routing. -
Production systems route because models are not interchangeable. They differ in capability, quality, context window, latency, price, modality, regional availability, rate limits, and reliability. Always using the strongest model wastes budget and latency; always using the cheapest model silently degrades the hard cases.
-
Routing does not require an LLM to choose an LLM. Deterministic rules are often the correct production design when tasks and requirements are known. A classifier or dynamic router is an optional later stage, not a maturity badge.
-
Routing becomes more sophisticated at scale because more signals appear: tenant policy, remaining deadline, provider health, measured quality, and cost ceilings. Those signals do not replace the original decision; they constrain and rank it.
-
Keep four ideas distinct. The AI Gateway is the provider-access boundary. Routing chooses among eligible models. Fallback is what happens when the preferred path cannot be used or fails. Evaluation tells you whether the chosen route is actually good enough.
Architecture Snapshot
Complexity
★★★☆☆
Audience
AI Engineers, Platform Engineers, Architects
Difficulty
Advanced
Typical Deployment
Application policy or gateway policy
Typical Latency
Keep the decision cheap vs generation
Scalability
Earned by mixed workloads and constraints
Availability Approach
Bounded fallback inside the request deadline
Read Time
~40 min
Last Updated
September 9, 2026
Recommended Starting Point
- Pin one model ID when the workload is uniform
- Add a task-to-model map when jobs actually differ
- Filter by hard constraints before optimizing cost or latency
- Treat fallback as a named, bounded alternative — not unlimited retries
Why This Matters
A prototype usually has one model ID in one handler. That pin is a routing policy with a single row. It is the right design for a long time.
The shape changes when the product has more than one kind of request. Classification does not need a long-context reasoning model. A vision request cannot go to a text-only endpoint. A tenant in a restricted region cannot use a provider that is not allowed. Interactive chat cannot wait on a batch-sized model. A billing incident appears when every lookup is sent to the most expensive tier “just in case.”
Those are not prompt problems. They are selection problems: given this request, which eligible model should run, why, and what happens if that choice is unavailable.
This guide answers a practical question: how does a production AI system decide which model should handle a request? It is not a ranking of models or gateway products. Production AI Stack is the capability map that lists routing as one role. AI Gateway is the provider-access boundary where that decision may execute. This page is only the decision itself.
The Problem Model Routing Solves
The production problem is mismatch between workload requirements and a single default model.
A system may have several models because they differ on dimensions that matter operationally:
| Dimension | Why it forces a choice |
|---|---|
| Capability | Tool calling, structured output, vision, reasoning, audio |
| Reasoning quality | Simple extraction versus high-stakes synthesis |
| Context window | A 200-token classifier versus a long document |
| Latency | Interactive UI versus overnight batch |
| Price | Unit cost per token or per request |
| Modality | Text, image, audio, or mixed inputs |
| Provider availability | Outages, capacity, and regional endpoints |
| Rate limits | RPM/TPM ceilings that make a preferred model unusable |
| Reliability | Error rates, timeout tails, and quality variance |
| Policy | Allowed providers, data-handling rules, contractual limits |
“Always use the strongest model” is often a poor production strategy. Frontier-quality generation is wasted on intent classification. It increases latency, spend, and rate-limit pressure without improving the product. It also concentrates failure: one provider incident becomes an outage for every task.
The opposite failure is “always use the cheapest model.” Cheap models are appropriate for many tasks. They are the wrong default for work that needs a longer context window, a required capability, or a measured quality bar. Cost-only routing looks efficient in a dashboard until the hard cases fail quietly.
The point is not that one model is universally best. The point is matching workload requirements to an appropriate model, then having a named alternative when that match cannot be executed.
| Without an explicit route | With a routing policy |
|---|---|
| Every handler hard-codes a model ID | Task class or request features select from a catalog |
| “The LLM” is treated as one interchangeable box | Capability and policy filter the eligible set |
| Cost and quality fight in ad hoc if/else | Hard constraints first, then ranked preferences |
| Provider errors become product outages | Bounded fallback under the remaining deadline |
| Nobody can answer why model B served this tenant | The decision emits a reason, policy version, and model ID |
Note
A single pinned model is still a route. Name it. The day a second model appears, you already have a table instead of a hunt through handlers.
How We Got Here
Teams did not start with a router. They started with one API call.
Diagram: From one model ID to a routing policy
timeline
title From one model ID to routing policy
Pin : One handler, one model ID
Tasks : Static map by job type
Providers : Capability and availability
Policy : Tenant, region, and eval
Routing starts as a pin. Extra sophistication is a response to mixed workloads and constraints, not a default platform layer.
| Era | What teams did | What started to hurt |
|---|---|---|
| One model ID | Call a provider SDK from the request handler | Almost nothing — this is still the right start |
| Several job types | Classification, summarization, and generation share one model | Latency and spend rise on work that did not need the large model |
| Several providers | Hosted APIs plus self-hosted or open models | Capability and failure modes diverge; IDs proliferate |
| Operational constraints | Cost ceilings, SLOs, tenant allow-lists | Ad hoc if/else becomes unexplainable |
| Measured routing | Golden sets and traces per route | Intuition about “better models” gets replaced by evidence |
Gateway libraries and multi-provider proxies packaged adapters and, often, a routing table. That packaging is convenient. It is not the definition of routing. A twenty-line if task == … map in application code is routing. A purchased “LLM router” with no policy and no eval is just another hop that still has to choose a model ID.
The useful lesson is the same one as the rest of the production AI stack: name the decision, then decide how thin the implementation can be.
What Is Model Routing?
Model routing is the decision process that determines which model or provider should handle a request.
Three parts of that definition matter:
- Decision process — it produces a choice (and, in production, a reason). It is not the model call itself.
- Which model or provider — the output is an eligible backend, not a rewritten prompt or a retrieval plan.
- Should handle a request — the input is one request with known or inferred requirements, not a global “best model” ranking.
It is not automatically:
- An LLM that chooses another LLM
- A classifier on every request
- A requirement to use multiple providers
- A centralized platform service
- A dynamic policy that changes per token
Those are optional implementations. The concept is the choice.
Nearby ideas that are not the same thing
| Term | What it is | How it relates to routing |
|---|---|---|
| Model selection | Choosing a model for a product, experiment, or eval campaign | Offline or design-time. Routing reuses that catalog per request |
| Model routing | Choosing which eligible model handles this request | Online decision |
| Model fallback | An alternative path when the preferred route cannot be used or fails | Recovery, not the initial choice |
| Load balancing | Spreading traffic across equivalent replicas of the same logical model | Capacity, not capability matching |
| Model orchestration | Sequencing retrieval, tools, agents, and generation | Control flow. Routing is one lookup inside that flow |
Selection answers “which models belong in the system at all?” Routing answers “which of those models should serve this request?” Fallback answers “what do we do when the preferred answer cannot run?” Load balancing assumes the candidates are equivalent. Orchestration decides what work to do; routing decides which model does a generation step.
Remember
Routing is a decision. Fallback is generally what happens when the preferred path cannot be used or fails. A production system may implement both in one component. The concepts should still be named separately, or incidents become un-debuggable.
How Model Routing Works
The simplest production router is a lookup.
if task == "classification":
use Model A
if task == "summarization":
use Model B
if task == "complex_reasoning":
use Model C
That is already model routing. No classifier. No second LLM. No multi-provider mesh. The system has a policy that maps a known task to a known model.
Why this counts: the request is not sent to “the default LLM.” It is sent to a model chosen because the task’s requirements match that model’s role. The policy can be a config file, a dictionary, or a gateway route table. The mechanism is boring on purpose.
Task-based routing
The request carries, or is assigned, a task class: classification, extraction, summarization, coding, reasoning, generation, rewrite, embedding, rerank, and so on. Each class has a default model. This is the usual starting point because product surfaces already know what they are asking for. A classification worker does not need to infer that it is doing classification.
Capability-based routing
Some requests need a capability the default model does not have: vision input, a context window large enough for the prompt, structured output, tool calling, or multilingual support. The router first asks “which models can do this at all?” and only then which of those is preferred. A text-only model is not a cheaper substitute for a vision request; it is an ineligible candidate.
Modality-based routing
Modality is a hard capability filter that is easy to forget in text-centric systems. Image, audio, and mixed inputs must go to models that accept those payloads. Treating modality as a soft preference produces failed calls, not cheaper ones.
These three forms are usually enough for an application with known jobs. Dynamic routing appears later, if the task class itself is unknown or the tradeoff between models is not stable.
Architecture
In a production AI system, routing sits on the path from application to model. It often runs inside an AI gateway, but it can live in the application or orchestrator instead. AI System Architecture treats orchestration as the control plane that decides what to ask; routing decides which model is asked.
Diagram: Architecture snapshot
flowchart TB
App[Application]
GW[AI Gateway / Provider Boundary]
Policy[Routing Policy]
Cands[Model / Provider Candidates]
Sel[Selected Model]
Resp[Response]
Signals[Routing signals]
App --> GW
GW --> Policy
Signals --> Policy
Policy --> Cands
Cands --> Sel
Sel --> Resp
The application reaches providers through a boundary. Routing policy selects among eligible candidates using request and operational signals.
The snapshot is intentionally thin. Routing is not a retrieval stack and not an agent loop. It consumes signals, produces a decision, and leaves execution — including retries and fallback — to a bounded caller.
Typical signals, covered in detail below:
| Signal | Role in the decision |
|---|---|
| Task | What kind of work this is |
| Capability | What the model must be able to do |
| Quality requirement | How wrong a bad answer is allowed to be |
| Latency | Interactive versus asynchronous budgets |
| Cost | Unit cost and remaining spend envelope |
| Availability | Health, capacity, rate limits |
| Tenant / policy | Allowed providers, regions, data-handling rules |
| Request characteristics | Token size, language, modality, known complexity |
| Evaluation results | Measured quality, latency, cost, and failure rates |
Not every system needs every signal. A single-product classifier with one provider needs a pin. A multi-tenant platform with several models needs most of the table.
Diagram: Request path through routing and execution
sequenceDiagram
participant App as Application
participant GW as Gateway
participant R as Routing policy
participant M as Selected model
App->>GW: Request, tenant, deadline
GW->>R: Catalog and signals
R-->>GW: Model ID and reason
GW->>M: Provider call
alt Success
M-->>GW: Response
GW-->>App: Response and route metadata
else Preferred path fails
GW->>R: Fallback in remaining deadline
R-->>GW: Alternate eligible model
end
Routing chooses. Execution may then fall back. Both steps must fit the caller’s remaining deadline.
Routing Signals
A production router is only as good as the inputs it is allowed to use. Start with the signals you already know. Add others when a named failure appears.
Task
Task is the most common signal because it is often already explicit. A summarization job, a support classifier, a coding copilot, and a long-form report generator are different products that happen to share an inference stack.
If the application does not know the task, someone still has to assign one: a route parameter, a workflow step name, or — later — a classifier. Guessing the task inside the frontier model you were trying to avoid is circular.
Capability
Capability filters are usually hard. Required examples:
- Context length sufficient for prompt plus expected output — see Context Windows
- Structured output or JSON schema support — see Structured Outputs
- Tool / function calling — see Function Calling and Tool Calling
- Vision or other modalities
- Reasoning depth the product has actually measured
- Language coverage
A model that cannot satisfy a required capability is not a cheaper option. It is not in the candidate set.
Quality requirement
Not every request has the same cost of being wrong. Tag extraction for an internal search index is not the same as a customer-facing legal summary. Quality is not “use the largest model.” It is a product requirement: what failure looks like, and whether a cheaper model meets the bar on a representative eval set.
Latency
Interactive paths have a deadline the user can feel. Batch paths can wait. A router that ignores latency will send UI traffic to a slow reasoning model because it scores higher on a quality benchmark. Pair this signal with Latency Optimization: the routing decision itself should be cheap compared with generation.
Cost
Budget-aware routing and cost ceilings are legitimate signals. They are not the only signals. Cost Optimization is mostly smaller prompts, caching, and sending work to an adequate model — not a gateway tax. A cost ceiling that can override a required capability or a tenant restriction is a policy bug.
Availability
Provider and model health, capacity, and rate limits are operational signals. Clearly unhealthy or unavailable candidates can be excluded from the eligible set. Capacity and load can also influence ranking among those that remain. A health check does not guarantee the provider will stay healthy for the actual request. Availability should not silently rewrite a quality-sensitive route without recording that a fallback ran.
Tenant or policy constraints
Allowed providers, regions, data-handling requirements, and contractual restrictions are hard constraints. They are closer to AI Security than to model preference. A router that can pick a disallowed provider because it is cheaper has bypassed the security boundary.
Request characteristics
Token size, language, modality, and known complexity (document length, number of tools, schema strictness) are features of this request. They often determine context-window eligibility and whether a “small” model can physically accept the payload.
Evaluation results
Offline and periodic eval can attach quality, latency, cost, and failure-rate priors to a route. That is how the system answers “is Model A actually better than Model B for this workload?” without relying on launch-week anecdotes. See Routing with Evaluation.
Engineering Insight
Unused signals are not a defect. A router with three reliable inputs and an explainable policy outperforms a router that consumes twenty noisy features.
Routing Strategies
Strategies range from a pin to a multi-stage policy. Later strategies are not automatically better.
A. Static / pinned routing
A known task always uses a known model. support_classifier → Model A. Version the pin. Log the model ID. This is the correct design when the workload is uniform.
B. Rule-based routing
Explicit policies determine the model: if/else, tables, or config. Rules can combine task, payload size, and tenant. They remain inspectable. Most production routers should start and often stay here.
C. Capability-based routing
Choose from models that satisfy required capabilities, then apply preferences. The catalog needs an honest capability matrix. If the matrix is wrong, the router will be confidently wrong.
D. Cost/latency-aware routing
Among acceptable models, prefer lower cost or lower latency. “Acceptable” is the important word. This strategy is a ranker on an already-filtered set, not a replacement for capability and policy filters.
E. Quality-aware routing
Use measured quality — golden-set scores, human review, or online outcome signals — to influence selection. Quality-aware routing without evaluation is just preference dressed as evidence.
F. Dynamic / classifier-based routing
A separate classifier or lightweight model estimates which route is appropriate, typically when the task class is not known up front or when complexity varies inside one product surface. This can be useful. It is not inherently superior. The classifier is another model to evaluate, version, and fail. If a request already knows it is a summarization job, do not classify it.
A learned router can go further: it uses historical preference or evaluation data to predict which model is likely to offer the best quality/cost trade-off for this request. RouteLLM is one example of that approach. It remains optional, and it still has to respect hard constraints.
G. Multi-stage routing
First filter by hard constraints, then rank remaining candidates on soft preferences. This is a natural production default once more than one signal exists. It is still deterministic if the filters and the scoring function are deterministic.
| Strategy | When it is enough | Extra failure mode |
|---|---|---|
| Pinned | One workload, one adequate model | Hidden when a second workload appears |
| Rule-based | Known tasks and requirements | Rule sprawl if every exception becomes a clause |
| Capability-based | Mixed modalities or APIs | Stale capability metadata |
| Cost/latency-aware | Eligible set already quality-safe | Optimizing the wrong objective |
| Quality-aware | You can measure the workload | Overfitting a tiny eval set |
| Classifier-based | Task is genuinely unknown | Router errors send work to the wrong tier |
| Multi-stage | Hard constraints plus preferences coexist | Complexity if stages are not named |
Important
Deterministic routing is often preferable when requirements are known and predictable. Dynamic or LLM-based routing is a tool for residual uncertainty, not a replacement for a policy you can already write down.
Hard Constraints vs Soft Preferences
Production routing policies almost always contain two different kinds of rule. Mixing them is a common source of security and quality incidents.
Hard constraints eliminate candidates. A violation means “this model must not serve this request,” regardless of cost or speed.
Examples:
- Provider not allowed for this tenant
- Required capability missing (vision, tools, structured output)
- Region or data-residency restriction
- Context window insufficient for the payload
- Tenant policy or contractual restriction
Soft preferences rank the candidates that remain.
Examples:
- Lower cost
- Lower latency
- Preferred provider
- Higher measured quality
- Better historical performance on this task
A router should eliminate candidates that violate hard constraints before optimizing soft preferences. If the remaining set is empty, fail closed or use an explicit, policy-approved fallback — do not “relax” a tenant restriction to save the request.
Simple example:
- Request: summarization, tenant
acme, ~8k input tokens, interactive SLO. - Hard filter: drop providers
acmecannot use; drop models with context window below the prompt; drop models without the summarization capability you actually require. - Remaining: Model B and Model C.
- Soft rank: Model B meets the latency SLO at lower cost; Model C scores slightly higher on the summarization eval.
- Policy says latency SLO beats marginal eval gain for this product surface → select Model B.
- If Model B is unavailable, fallback may try Model C if it still satisfies the hard filters.
The same request must not select a disallowed provider because it is 10% cheaper. That is not ranking. That is a constraint failure.
Step-by-Step Flow
A typical production decision looks like this:
- Receive the request with tenant identity, payload, remaining deadline, and any explicit task class.
- Identify the task from a trusted field, workflow step, or — only if needed — a classifier.
- Apply tenant and provider restrictions so disallowed backends never enter ranking.
- Filter by required capabilities and context-window fit.
- Remove clearly unhealthy or unavailable candidates using health signals and current rate-limit state.
- Rank remaining candidates by the organization’s preference order (quality, latency, cost, preferred provider, and capacity where it is known).
- Select a model and record the reason plus policy version.
- Execute the request with a timeout inside the remaining deadline.
- Observe the result — success, error class, tokens, latency — and, if needed, run a bounded fallback.
If step 5 or 6 empties the set, the system should fail with an explicit “no eligible model” outcome rather than widening hard constraints.
Production Control-Flow Example
The following is illustrative pseudocode, not production-ready code and not a vendor SDK. It shows the control flow: filter, then rank, then execute with a separate fallback path.
# Illustrative pseudocode — not a production SDK and not a real vendor API.
def identify_task(request):
# Routing-relevant task identity comes from trusted application context,
# not arbitrary user input.
return request.trusted_task
def eligible_models(request, catalog, policy, health):
task = identify_task(request)
candidates = catalog.for_task(task)
allowed = []
for model in candidates:
if not policy.allows(request.tenant, model):
continue
if not model.satisfies(request.required_capabilities):
continue
# Full context need: input + expected output + tool/schema overhead.
if model.context_window < request.required_context_tokens:
continue
if not health.is_available(model):
continue
allowed.append(model)
return allowed
def route(request, catalog, policy, health):
candidates = eligible_models(request, catalog, policy, health)
if not candidates:
raise NoEligibleModel("no model satisfied hard constraints")
ranked = sorted(
candidates,
key=lambda model: policy.preference_key(model, request, health),
)
selected = ranked[0]
return RouteDecision(
model=selected,
reason=policy.explain(selected, request),
remaining=ranked[1:],
)
def execute_with_fallback(request, decision, caller_deadline):
attempts = [decision.model, *decision.remaining[:1]] # bound the chain
last_error = None
for model in attempts:
remaining = caller_deadline.remaining()
if remaining <= 0:
break
try:
return call_model(model, request, timeout=remaining)
except TransientProviderError as err:
last_error = err
# Same-backend retry (not shown): honor Retry-After only if the
# remaining deadline still allows that wait. Moving to the next
# eligible backend is a separate decision and must not inherit
# this provider's Retry-After delay.
if not should_fallback(err):
raise
continue
raise RouteExhausted(last_error)
What the sketch is trying to make obvious:
- Task identity comes from trusted application context, not from arbitrary user input or an extra LLM by default.
- Tenant policy and capability checks happen before ranking.
- Health should generally gate eligibility; capacity and operational signals can also influence ranking among eligible candidates. A health check does not guarantee the provider will remain healthy for the actual request.
- The reason is part of the decision, not a log line added later.
- Fallback is a short, explicit list under the remaining deadline — not an unbounded retry loop. Same-backend retries may honor
Retry-Afterwhen the deadline allows; fallback to another backend does not inherit that delay.
Before using a fallback candidate, production implementations should re-check dynamic eligibility signals such as current availability and rate-limit state; static policy and capability constraints remain part of the original decision.
Wire this behind whatever adapter you already use to call providers. If that adapter is a shared boundary, it is an AI Gateway. The router does not need to own credentials, HTTP retries, or token accounting.
Advanced Routing: Quality, Cost, and Latency
Once more than one eligible model exists, routing becomes a policy / optimization problem rather than a lookup table.
A router may be trying to improve several dimensions at once:
| Objective | Typical pressure |
|---|---|
| Quality | Prefer models that score higher on the task |
| Latency | Prefer models that meet the interactive SLO |
| Cost | Prefer lower unit cost within the envelope |
| Availability | Prefer backends that can accept the request now |
These objectives conflict.
- Highest quality often costs more and may be slower.
- Lowest cost may increase latency or miss a quality bar.
- Lowest latency shrinks the candidate set (some models cannot meet the SLO).
- Provider availability can change the optimal route minute to minute.
There is no universal scoring function. An interactive support classifier may rank latency, cost, quality. A high-risk summary may rank quality, policy, latency, cost. The same catalog can serve both products if the policy object is per product or per tenant, not a global “best model” score.
Keep the math modest. A lexicographic order (satisfy SLO, then minimize cost) or a short weighted score on an already-filtered set is enough for most teams. Academic multi-objective optimization is optional. Explainability is not: if operators cannot say why Model B won, you cannot debug a regression.
Routing with Evaluation
Routing policies should be informed by evaluation, not by intuition alone.
A model that “feels smarter” is not a route. Compare candidates on representative workloads for each task class:
- Quality by task (correctness, format, faithfulness where retrieval is involved)
- Latency (p50/p95, not only mean)
- Cost per successful request
- Failure rates (timeouts, refusals, schema violations, provider errors)
Evaluation answers a question routing cannot answer by construction: is Model A actually better for this workload than Model B? Public leaderboards and benchmarks shortlist models. Product eval decides whether a route is safe to ship. See also LLM Evaluation for generation-level measurement and Agent Evaluation when the “model” sits inside a tool loop — a cheaper model that calls tools well may outperform a larger model that does not.
Practical consequences:
- Pair every new route with a slice of the golden set for that task.
- Re-evaluate when providers change versions, prices, or behavior.
- Do not promote a cheap model on cost dashboards unless quality on the hard slice is still acceptable.
- Traffic splitting and canaries are routing operations; they need the same eval hooks as a full cutover.
Warning
A routing change is a behavior change. Treat it like a deploy: version the policy, canary it, and keep a rollback to the previous pin.
Routing and Observability
Routing without visibility becomes folklore. You do not need a separate observability architecture on this page — see Observability for traces, metrics, and logs as a system. You do need the route to be a first-class field on the request span.
Minimum useful record:
| Field | Why it exists |
|---|---|
| Selected model ID | What actually ran |
| Provider | Where it ran |
| Routing reason / policy ID | Why this candidate won |
| Latency | Decision time versus generation time |
| Token usage and cost | Attribution per tenant and route |
| Error class | Timeout, 429, 5xx, capability miss, policy deny |
| Retries / fallbacks | Whether the preferred path held |
| Outcome / eval signals | Online quality hooks where you have them |
If you cannot answer “why was this model selected for this tenant last Thursday?”, you do not yet have a routing policy. You have a default that drifted.
Prompts, completions, tool arguments, and retrieved content should not be captured automatically merely because routing is observable; they may contain sensitive data and should follow the application’s telemetry redaction and retention policy.
Keep the routing decision itself cheap and synchronous enough to trace. A router that calls a slow LLM to choose a fast LLM has already spent the latency budget it was created to save.
Design Decisions
| Question | Simpler choice | More advanced choice | When to prefer the advanced choice |
|---|---|---|---|
| One model or many? | Pin one ID | Catalog of task-specific models | Workloads differ in capability, quality, or cost |
| Static rules or dynamic routing? | Config / if-else | Classifier or scored router | Task or complexity is not known from the application |
| Single provider or multi-provider? | One vendor SDK | Provider + model catalog | Availability, region, or capability gaps are real |
| Cost vs quality optimization? | Meet a quality bar, then cut cost | Multi-objective ranker | You have eval coverage and conflicting SLOs |
| Where does routing live? | Application-local function | Shared policy on the AI gateway | Several services would otherwise copy the same table |
| Fallback depth? | Fail the request | One named alternate | Availability matters more than identical answers |
| How is the decision explained? | Log model ID | Policy version + reason code | You operate more than one route in production |
More complexity is not automatically better. Each extra signal is another way to be wrong and another field to keep fresh.
Common routing patterns
These are named ways to fill the table above. Use them when the situation matches, not as a checklist.
| Pattern | What it does | When it makes sense |
|---|---|---|
| Cheap-first | Prefer the lowest-cost eligible model | Quality bar is met by the cheap model on eval |
| Quality-first | Prefer the best measured model, ignore extra unit cost | Cost of errors dominates token cost |
| Latency-sensitive | Drop candidates that miss the SLO before ranking | Interactive UI or strict p95 |
| Capability routing | Filter by vision, tools, context, schema, language | Mixed modalities or APIs |
| Tenant-aware | Allow-lists and data-handling rules first | Multi-tenant or regulated products |
| Regional / provider routing | Choose a backend that may serve this region | Residency, sovereignty, or regional capacity |
| Fallback routing | Alternate eligible model after a classified failure | Provider blips without relaxing hard constraints |
| Traffic splitting | Send a percentage to a candidate for eval | Comparing models with live traffic |
| Gradual migration | Canary from Model A to Model B by task or tenant | Replacing a pin without a big-bang cutover |
Cheap-first without a quality gate is how hard cases rot. Quality-first without a cost envelope is how a classifier becomes a frontier-model bill. Traffic splitting without assignment in traces makes the experiment unreadable.
Comparisons
Model routing vs model selection
Selection is catalog design: which models are offered, at which versions, for which tasks. Routing is the per-request use of that catalog. You can select models quarterly and still route every request. Confusing the two leads to “we picked Claude for the company” as if that were a request policy.
Model routing vs fallback
Routing: “Use Model B for summarization because it meets the task requirements.”
Fallback: “Model B failed or became unavailable, so try Model C.”
Fallback can be implemented next to routing operationally. It is still a different question. Fallback is triggered by timeout, provider errors, rate limits, or temporary unavailability, and it must fit the remaining request deadline. It is not an unlimited retry chain. If the preferred backend returns Retry-After, a same-backend retry should honor that delay only when the remaining deadline allows it. Switching to another eligible backend is a separate decision and should not blindly inherit the original provider’s Retry-After. A cheaper or different model is a behavior change, not a hot spare with the same answers. Pair fallbacks with evaluation if quality matters — the same warning as in the AI Gateway guide.
Diagram: Preferred route versus fallback
stateDiagram-v2
[*] --> Route
Route --> Execute: Selected model
Execute --> Done: Success
Execute --> Fallback: Timeout, rate limit, or unavailable
Fallback --> Execute: Alternate eligible model
Fallback --> Fail: No remaining budget or candidates
Done --> [*]
Fail --> [*]
Fallback starts after the preferred route is known. It does not replace the initial decision, and it must stop.
Model routing vs load balancing
Load balancing spreads traffic across equivalent replicas (same logical model, multiple instances or keys). Routing chooses among non-equivalent models based on request requirements. After a model is selected, load balancing may still distribute traffic across equivalent replicas of that model. Using a load balancer as a capability router will send vision requests to text replicas whenever those replicas are idle.
Model routing vs orchestration
Orchestration sequences steps: retrieve, call tools, generate, validate. Routing chooses the model for a generation step. An agent framework can contain a router; it is not a router by itself. See AI Agents and AI System Architecture.
Provider routing vs model routing
Model routing asks: which model should handle the request?
Provider routing asks: which provider should serve it?
They combine. A logical requirement such as “interactive summarization, 32k context, tenant-allowed” may have several provider-specific model IDs. Separating the logical capability from the provider implementation is useful: you can fail over providers without changing the product’s idea of the task, and you can change a model ID without pretending the task changed.
A gateway often performs provider adapters; the router decides which adapter and which ID. See AI Gateway for that boundary. Do not duplicate credential, authentication, rate-limit, or adapter design here.
Model routing and the AI Gateway
Conceptual boundary:
Application → AI Gateway → routing / policy → provider / model
The gateway can provide the provider-access and policy enforcement point. Routing determines which eligible model or provider should handle the request. Either can exist without the other: a single-provider app can route between tiers in application code with no gateway; a gateway can exist solely for credentials, timeouts, and traces with one model behind it.
Details of credentials, caller authentication, rate limiting, observability pipelines, and provider adapters belong in AI Gateway. This guide only needs the split: boundary versus decision.
Routing Policy and Precedence
Conflicting requirements need an explicit order. One common — not universal — precedence is:
- Security / tenant restrictions
- Required capability
- Context-window requirements
- Reliability / availability
- Quality target
- Latency target
- Cost preference
Organizations define their own order. A latency-critical classifier may place SLO above measured quality. A regulated workload may place region above availability (better to fail than to leave the region). The key concept is that routing should be explainable: the system should be able to answer “why was this model selected?”
Store precedence in the policy, not in tribal knowledge. When two teams disagree about cost versus quality, that is a product decision to encode, not a reason to add another classifier.
Common Mistakes
-
Routing based only on model price. The cheap model that cannot see the image, fill the schema, or meet the quality bar is not a saving.
-
Assuming the largest model is always best. Extra reasoning depth does not help a two-class intent label. It does add latency, spend, and rate-limit risk.
-
Using an LLM router when deterministic rules are sufficient. If the application already knows the task, a map is simpler, faster, and easier to evaluate.
-
Allowing routing to bypass tenant or security policy. Soft preferences must not resurrect a candidate the hard filter removed.
-
Treating fallback as unlimited retries. Completions are often billable per attempt. Cap the chain, stay inside the caller deadline, and honor
Retry-Afteronly when retrying the same backend if that wait still fits. -
Ignoring context-window limits. Sending an over-long prompt to a small model is not routing. It is a guaranteed failure that should have been filtered.
-
Routing without measuring outcomes. A policy that is never eval’d will drift as providers change versions underneath stable IDs — or as IDs are swapped without a golden set.
-
Creating overly complicated routing rules. Every special case is a branch that will not be tested. Prefer a small table plus an explicit exception list.
-
Making routing decisions impossible to explain. If the only artifact is “the router chose B,” you cannot tell a policy bug from a provider outage.
-
Changing routing policies without regression evaluation. A canary that does not include the hard slice of the task will look green.
Where It Breaks Down
Poor task classification. If the task label is wrong, every downstream filter is applied to the wrong catalog slice. Trusted application labels beat inferred labels.
Insufficient evaluation data. Quality-aware and cheap-first policies both need a representative set. A dozen happy-path prompts will not detect the long-document failure.
Rapidly changing provider and model behavior. Versionless aliases move under you. Pin IDs, re-eval on change, and treat silent alias updates as deploys.
Correlated provider failures. Multi-provider fallback does not help if both providers share a region, a GPU shortage, or the same upstream dependency. Diversity on paper is not diversity in the incident.
Hidden quality regressions. A cheaper route can pass format checks and fail usefulness. Online feedback loops that only measure thumbs-up will hide task-specific damage.
Routing complexity. A policy no one can simulate in staging will not be debugged in production. Complexity is a reliability hazard.
Stale routing policies. Capability matrices, prices, and context windows go stale. A quarterly catalog review is part of running the router.
Feedback loops. If online quality signals train the classifier that chooses the route, a biased sample can lock traffic onto a worse model. Keep a held-out eval path.
Unpredictable workloads. If every request is a new shape, static routes underfit and classifiers overfit. You may want a stronger default model and fewer routes, not more machinery.
Routing is only as good as the signals and measurements behind it. When those are weak, pin fewer models.
When NOT to Use Model Routing
A single-model application is often the correct architecture.
Skip a multi-model router when:
- One model already meets quality, latency, and cost for the actual workload
- You cannot name how tasks differ in a way a second model would improve
- You do not yet measure quality, so a second route would be an untested hypothesis
- The “router” would exist only because an architecture diagram included one
- You would need an LLM to classify work that the application already labels
Use routing when several of these are true:
- Multiple models genuinely provide different value (capability, quality, or cost)
- Workloads differ materially by task, modality, or risk
- Cost, latency, or availability tradeoffs are real and measurable
- Provider diversity is an operational requirement you can test
- Migration or experimentation needs controlled selection (canary, split, rollback)
Do not introduce routing merely because it sounds architecturally sophisticated. An unmeasured mesh of models is harder to operate than a pinned ID with traces.
Decision tree: do you need more than a pin?
flowchart TD
Start[More than one model or provider?] -->|No| Pin[Pin one model ID]
Start -->|Yes| Diff{Do tasks differ in capability, quality, latency, or cost?}
Diff -->|No| Avail{Need availability failover?}
Avail -->|No| Pin
Avail -->|Yes| Fallback[Named fallback]
Diff -->|Yes| Known{Are requirements known?}
Known -->|Yes| Rules[Deterministic routing]
Known -->|No| Measure[Evaluate then add rules]
Measure --> Still{Still unpredictable?}
Still -->|No| Rules
Still -->|Yes| Class[Optional classifier]
Start from a pin. Add rules when jobs differ. Add a classifier only when the task is not already known.
Running in Production
Best Practice
Keep decision logic bounded and deterministic where possible. Filter hard constraints first. Cap fallbacks. Version the policy. Emit the reason. Re-evaluate routes on a schedule, not only after incidents.
| Dimension | Guidance |
|---|---|
| Decision logic | Bounded, testable functions or tables — not an unbounded agent loop |
| Determinism | Prefer rules you can replay in staging with the same inputs |
| Deadlines | Routing + generation + fallback must fit the caller’s remaining budget |
| Health | Use availability signals; do not wait forever to discover a 429 |
| Fallback | Named, short, constraint-preserving; honor Retry-After on same-backend retries when the remaining deadline allows — do not inherit that delay onto another backend |
| Observability | Model, provider, policy version, reason, tokens, errors, fallback flag |
| Evaluation | Golden set per task class; block policy deploys on regressions |
| Versioning | Policy IDs like prompt and index versions |
| Rollout | Canary by task or tenant; keep an instant rollback to the previous pin |
| Security | Tenant restrictions are not optional ranker features |
Production checklist
- Every request records
model_id, provider, and policy version - Hard constraints (tenant, region, capability, context window) run before ranking
- Soft preferences cannot resurrect a disallowed candidate
- Timeouts on every provider call; fallback stays inside the remaining deadline
- Fallback chain is capped and logged; same-backend retries honor
Retry-Afterwhen the remaining deadline allows it; fallback to another backend does not inherit that delay - Each route has a representative eval slice; policy changes go through it
- Health / rate-limit signals can mark a model temporarily ineligible
- Operators can answer “why this model?” from traces
- A single-model pin remains a supported configuration
- Alias-free model IDs, or aliases treated as deploys when they move
Interview Questions
-
What is model routing?
The decision process that chooses which eligible model or provider should handle a request. It can be a static map. It does not require an LLM router. -
How is routing different from fallback?
Routing is the preferred choice given current signals. Fallback is an alternative path when that choice cannot be used or fails. Fallback must be bounded. -
When would deterministic routing be better than an LLM router?
When the application already knows the task and requirements. Rules are faster, cheaper, testable, and explainable. Use a classifier when the task is genuinely unknown. -
How would you route between models with different cost, latency, and quality?
Filter by hard constraints first. Rank the remainder with an explicit precedence (for example: meet quality bar, meet SLO, then minimize cost). Measure all three on a golden set. -
How would you enforce tenant restrictions?
As a hard filter before ranking, using authenticated tenant identity. Never as a soft preference the ranker can override. -
How would you know whether a routing policy is actually working?
Per-route eval (quality, latency, cost, errors), traces that include the reason, and regression gates on policy changes. Dashboards that only show spend are not sufficient. -
Where should routing live relative to an AI gateway?
The gateway is the provider-access boundary. Routing is a policy that may run there or in the application. Either can exist without the other. See AI Gateway. -
How would you safely change routing rules in production?
Version the policy, canary by task or tenant, compare against the previous pin on the golden set, and keep an instant rollback. Treat provider alias changes the same way.
Related Guides
This guide is the decision process. Production AI Stack is the capability map. AI Gateway is the provider-access boundary. AI System Architecture is the platform blueprint.
Architecture:
- Production AI Stack — where routing sits among other capabilities
- AI Gateway — adapter, credentials, and policy enforcement point
- AI System Architecture — orchestration as control plane
- AI Copilot Architecture — application pattern that may route each turn
- Enterprise RAG Architecture — workload-based generation after retrieval
Foundations:
- Large Language Models · Tokens · Context Windows · Function Calling · Structured Outputs · Decision Models (cheap, calibrated route classification)
Operations:
- Evaluation · LLM Evaluation · Agent Evaluation · Observability · Cost Optimization · Latency Optimization · AI Security · Caching
Agents (routing is not orchestration):
Rankings: Best AI APIs — hosted access options, not a mandate to route across all of them
Tools: LiteLLM — one implementation of adapters and routing tables, not a required architecture
Learning path: Become an AI Engineer
Diagram: Recommended reading around this guide
flowchart LR
Stack[Production AI Stack]
GW[AI Gateway]
Route[Model Routing]
Eval[Evaluation]
Stack --> GW
Stack --> Route
GW --> Route
Route --> Eval
Read the stack map for composition, the gateway guide for the access boundary, and evaluation for whether a route is actually good.
Learning Path
Prerequisites: Large Language Models
Next topics: AI Gateway · Production AI Stack · AI Copilot Architecture · Evaluation · Cost Optimization · Observability
Estimated time: 40 min · Difficulty: Advanced
FAQs
Is model routing the same as model selection?
No. Selection is deciding which models belong in the catalog (design-time or eval-time). Routing is deciding which eligible catalog entry handles this request. You need both; they are not the same job.
Does model routing require an LLM?
No. A task-to-model map is routing. An LLM or classifier router is optional when the task is not already known. It is not the definition.
Is model routing the same as fallback?
No. Routing chooses the preferred eligible model. Fallback is used when that path cannot be used or fails. A system may implement both together; the questions remain different.
Should routing happen inside the AI gateway?
It can. The gateway is a natural enforcement point when several services share policy. Application-local routing is correct when one service owns the workloads. See AI Gateway.
How do you choose between cost and quality?
Encode a bar, not a vibe. Filter to models that meet the quality requirement on eval, then optimize cost (or latency) among those that remain. If no model meets the bar at the allowed cost, that is a product decision — not a ranker tie-break.
When should a system use multiple providers?
When a second provider supplies a capability, region, or availability property the first cannot, and you can test failover as a behavior change. Multiple providers are not required for routing between tiers of one provider.
How often should routing policies be reevaluated?
Whenever models, prices, context windows, or quality bars change — and on a regular cadence even if nothing was announced. Alias-based model IDs can move without a deploy.
Is load balancing a substitute for routing?
No. Load balancing assumes equivalent backends. Routing assumes they are not equivalent. You may load-balance replicas of the model you just selected.
Can routing live in the application instead of a platform service?
Yes. A function that returns a model ID from a table is a router. Extract a shared policy when several codebases copy the same rules.
What if no candidate survives hard constraints?
Fail closed, or use a pre-approved degraded path that still satisfies policy (for example a refusal message). Do not relax tenant or residency rules to obtain a completion.
Should every request go through a complexity classifier?
Only if complexity is not already known and eval shows the classifier improves outcomes enough to pay for its errors and latency. Many systems already know the job type.
References
- OpenAI — Models and pricing
- Anthropic — Models
- Google AI — Gemini models
- HTTP Semantics — Retry-After
- OpenTelemetry Documentation
- LiteLLM — Routing
- Ong et al., RouteLLM: Learning to Route LLMs with Preference Data
Further Reading
- Production AI Stack
- AI Gateway
- AI System Architecture
- Evaluation
- Cost Optimization
- Observability
- Best AI APIs
Key Takeaways
- Model routing is the per-request decision that chooses an eligible model or provider — a static task map already counts.
- It exists because models differ in capability, quality, latency, cost, availability, and policy; neither “always strongest” nor “always cheapest” is a production strategy.
- Hard constraints filter; soft preferences rank. Tenant and capability rules are not optional score terms.
- Routing ≠ fallback ≠ gateway ≠ evaluation. The gateway is a boundary; fallback is recovery; eval tells you whether the route works.
- Deterministic routing is often the right design; an LLM router is for residual uncertainty, not prestige.
- Observe model, provider, reason, and policy version, and re-evaluate routes when models or workloads change.
- A single pinned model is a valid architecture. Add routing when mixed workloads actually require it.