Architecture

Model Routing: Choosing the Right Model for Every Request Guide

How production AI systems route requests across models and providers using capability, quality, latency, cost, availability, and policy constraints.

40 min readAdvancedLast reviewed: 9 September 2026

Quick Summary

Model routing is the decision process that selects which eligible model or provider should handle a request, using task, capability, quality, latency, cost, availability, and policy signals.

One Analogy

Model routing is a dispatch desk — it assigns each job to a qualified worker. It is not another worker that debates who should do the job.

Engineering Rule

Eliminate candidates that violate hard constraints before optimizing for cost, latency, or quality. Do not introduce an LLM router when a deterministic policy already names the route.

Model Routing in One Sentence

Model Routing =

  • Task and capability match
  • Hard policy constraints
  • Quality / latency / cost preferences
  • Availability

Fallback is not part of that definition. It is a separate, bounded recovery path used when the preferred route cannot be used or fails.

TL;DR

  • Model routing is a decision, not a product. Given a request, the system chooses which eligible model or provider should handle it. A static map such as classification → small model, summarization → efficient model, complex reasoning → stronger model is already routing.

  • Production systems route because models are not interchangeable. They differ in capability, quality, context window, latency, price, modality, regional availability, rate limits, and reliability. Always using the strongest model wastes budget and latency; always using the cheapest model silently degrades the hard cases.

  • Routing does not require an LLM to choose an LLM. Deterministic rules are often the correct production design when tasks and requirements are known. A classifier or dynamic router is an optional later stage, not a maturity badge.

  • Routing becomes more sophisticated at scale because more signals appear: tenant policy, remaining deadline, provider health, measured quality, and cost ceilings. Those signals do not replace the original decision; they constrain and rank it.

  • Keep four ideas distinct. The AI Gateway is the provider-access boundary. Routing chooses among eligible models. Fallback is what happens when the preferred path cannot be used or fails. Evaluation tells you whether the chosen route is actually good enough.

Architecture Snapshot

Complexity

★★★☆☆

Audience

AI Engineers, Platform Engineers, Architects

Difficulty

Advanced

Typical Deployment

Application policy or gateway policy

Typical Latency

Keep the decision cheap vs generation

Scalability

Earned by mixed workloads and constraints

Availability Approach

Bounded fallback inside the request deadline

Read Time

~40 min

Last Updated

September 9, 2026

Recommended Starting Point

  • Pin one model ID when the workload is uniform
  • Add a task-to-model map when jobs actually differ
  • Filter by hard constraints before optimizing cost or latency
  • Treat fallback as a named, bounded alternative — not unlimited retries

Why This Matters

A prototype usually has one model ID in one handler. That pin is a routing policy with a single row. It is the right design for a long time.

The shape changes when the product has more than one kind of request. Classification does not need a long-context reasoning model. A vision request cannot go to a text-only endpoint. A tenant in a restricted region cannot use a provider that is not allowed. Interactive chat cannot wait on a batch-sized model. A billing incident appears when every lookup is sent to the most expensive tier “just in case.”

Those are not prompt problems. They are selection problems: given this request, which eligible model should run, why, and what happens if that choice is unavailable.

This guide answers a practical question: how does a production AI system decide which model should handle a request? It is not a ranking of models or gateway products. Production AI Stack is the capability map that lists routing as one role. AI Gateway is the provider-access boundary where that decision may execute. This page is only the decision itself.

The Problem Model Routing Solves

The production problem is mismatch between workload requirements and a single default model.

A system may have several models because they differ on dimensions that matter operationally:

Dimension Why it forces a choice
Capability Tool calling, structured output, vision, reasoning, audio
Reasoning quality Simple extraction versus high-stakes synthesis
Context window A 200-token classifier versus a long document
Latency Interactive UI versus overnight batch
Price Unit cost per token or per request
Modality Text, image, audio, or mixed inputs
Provider availability Outages, capacity, and regional endpoints
Rate limits RPM/TPM ceilings that make a preferred model unusable
Reliability Error rates, timeout tails, and quality variance
Policy Allowed providers, data-handling rules, contractual limits

“Always use the strongest model” is often a poor production strategy. Frontier-quality generation is wasted on intent classification. It increases latency, spend, and rate-limit pressure without improving the product. It also concentrates failure: one provider incident becomes an outage for every task.

The opposite failure is “always use the cheapest model.” Cheap models are appropriate for many tasks. They are the wrong default for work that needs a longer context window, a required capability, or a measured quality bar. Cost-only routing looks efficient in a dashboard until the hard cases fail quietly.

The point is not that one model is universally best. The point is matching workload requirements to an appropriate model, then having a named alternative when that match cannot be executed.

Without an explicit route With a routing policy
Every handler hard-codes a model ID Task class or request features select from a catalog
“The LLM” is treated as one interchangeable box Capability and policy filter the eligible set
Cost and quality fight in ad hoc if/else Hard constraints first, then ranked preferences
Provider errors become product outages Bounded fallback under the remaining deadline
Nobody can answer why model B served this tenant The decision emits a reason, policy version, and model ID

Note

A single pinned model is still a route. Name it. The day a second model appears, you already have a table instead of a hunt through handlers.

How We Got Here

Teams did not start with a router. They started with one API call.

Diagram: From one model ID to a routing policy

timeline
    title From one model ID to routing policy
    Pin : One handler, one model ID
    Tasks : Static map by job type
    Providers : Capability and availability
    Policy : Tenant, region, and eval

Routing starts as a pin. Extra sophistication is a response to mixed workloads and constraints, not a default platform layer.

Era What teams did What started to hurt
One model ID Call a provider SDK from the request handler Almost nothing — this is still the right start
Several job types Classification, summarization, and generation share one model Latency and spend rise on work that did not need the large model
Several providers Hosted APIs plus self-hosted or open models Capability and failure modes diverge; IDs proliferate
Operational constraints Cost ceilings, SLOs, tenant allow-lists Ad hoc if/else becomes unexplainable
Measured routing Golden sets and traces per route Intuition about “better models” gets replaced by evidence

Gateway libraries and multi-provider proxies packaged adapters and, often, a routing table. That packaging is convenient. It is not the definition of routing. A twenty-line if task == … map in application code is routing. A purchased “LLM router” with no policy and no eval is just another hop that still has to choose a model ID.

The useful lesson is the same one as the rest of the production AI stack: name the decision, then decide how thin the implementation can be.

What Is Model Routing?

Model routing is the decision process that determines which model or provider should handle a request.

Three parts of that definition matter:

  • Decision process — it produces a choice (and, in production, a reason). It is not the model call itself.
  • Which model or provider — the output is an eligible backend, not a rewritten prompt or a retrieval plan.
  • Should handle a request — the input is one request with known or inferred requirements, not a global “best model” ranking.

It is not automatically:

  • An LLM that chooses another LLM
  • A classifier on every request
  • A requirement to use multiple providers
  • A centralized platform service
  • A dynamic policy that changes per token

Those are optional implementations. The concept is the choice.

Nearby ideas that are not the same thing

Term What it is How it relates to routing
Model selection Choosing a model for a product, experiment, or eval campaign Offline or design-time. Routing reuses that catalog per request
Model routing Choosing which eligible model handles this request Online decision
Model fallback An alternative path when the preferred route cannot be used or fails Recovery, not the initial choice
Load balancing Spreading traffic across equivalent replicas of the same logical model Capacity, not capability matching
Model orchestration Sequencing retrieval, tools, agents, and generation Control flow. Routing is one lookup inside that flow

Selection answers “which models belong in the system at all?” Routing answers “which of those models should serve this request?” Fallback answers “what do we do when the preferred answer cannot run?” Load balancing assumes the candidates are equivalent. Orchestration decides what work to do; routing decides which model does a generation step.

Remember

Routing is a decision. Fallback is generally what happens when the preferred path cannot be used or fails. A production system may implement both in one component. The concepts should still be named separately, or incidents become un-debuggable.

How Model Routing Works

The simplest production router is a lookup.

if task == "classification":
    use Model A

if task == "summarization":
    use Model B

if task == "complex_reasoning":
    use Model C

That is already model routing. No classifier. No second LLM. No multi-provider mesh. The system has a policy that maps a known task to a known model.

Why this counts: the request is not sent to “the default LLM.” It is sent to a model chosen because the task’s requirements match that model’s role. The policy can be a config file, a dictionary, or a gateway route table. The mechanism is boring on purpose.

Task-based routing

The request carries, or is assigned, a task class: classification, extraction, summarization, coding, reasoning, generation, rewrite, embedding, rerank, and so on. Each class has a default model. This is the usual starting point because product surfaces already know what they are asking for. A classification worker does not need to infer that it is doing classification.

Capability-based routing

Some requests need a capability the default model does not have: vision input, a context window large enough for the prompt, structured output, tool calling, or multilingual support. The router first asks “which models can do this at all?” and only then which of those is preferred. A text-only model is not a cheaper substitute for a vision request; it is an ineligible candidate.

Modality-based routing

Modality is a hard capability filter that is easy to forget in text-centric systems. Image, audio, and mixed inputs must go to models that accept those payloads. Treating modality as a soft preference produces failed calls, not cheaper ones.

These three forms are usually enough for an application with known jobs. Dynamic routing appears later, if the task class itself is unknown or the tradeoff between models is not stable.

Architecture

In a production AI system, routing sits on the path from application to model. It often runs inside an AI gateway, but it can live in the application or orchestrator instead. AI System Architecture treats orchestration as the control plane that decides what to ask; routing decides which model is asked.

Diagram: Architecture snapshot

flowchart TB
    App[Application]
    GW[AI Gateway / Provider Boundary]
    Policy[Routing Policy]
    Cands[Model / Provider Candidates]
    Sel[Selected Model]
    Resp[Response]
    Signals[Routing signals]
    App --> GW
    GW --> Policy
    Signals --> Policy
    Policy --> Cands
    Cands --> Sel
    Sel --> Resp

The application reaches providers through a boundary. Routing policy selects among eligible candidates using request and operational signals.

The snapshot is intentionally thin. Routing is not a retrieval stack and not an agent loop. It consumes signals, produces a decision, and leaves execution — including retries and fallback — to a bounded caller.

Typical signals, covered in detail below:

Signal Role in the decision
Task What kind of work this is
Capability What the model must be able to do
Quality requirement How wrong a bad answer is allowed to be
Latency Interactive versus asynchronous budgets
Cost Unit cost and remaining spend envelope
Availability Health, capacity, rate limits
Tenant / policy Allowed providers, regions, data-handling rules
Request characteristics Token size, language, modality, known complexity
Evaluation results Measured quality, latency, cost, and failure rates

Not every system needs every signal. A single-product classifier with one provider needs a pin. A multi-tenant platform with several models needs most of the table.

Diagram: Request path through routing and execution

sequenceDiagram
    participant App as Application
    participant GW as Gateway
    participant R as Routing policy
    participant M as Selected model
    App->>GW: Request, tenant, deadline
    GW->>R: Catalog and signals
    R-->>GW: Model ID and reason
    GW->>M: Provider call
    alt Success
        M-->>GW: Response
        GW-->>App: Response and route metadata
    else Preferred path fails
        GW->>R: Fallback in remaining deadline
        R-->>GW: Alternate eligible model
    end

Routing chooses. Execution may then fall back. Both steps must fit the caller’s remaining deadline.

Routing Signals

A production router is only as good as the inputs it is allowed to use. Start with the signals you already know. Add others when a named failure appears.

Task

Task is the most common signal because it is often already explicit. A summarization job, a support classifier, a coding copilot, and a long-form report generator are different products that happen to share an inference stack.

If the application does not know the task, someone still has to assign one: a route parameter, a workflow step name, or — later — a classifier. Guessing the task inside the frontier model you were trying to avoid is circular.

Capability

Capability filters are usually hard. Required examples:

  • Context length sufficient for prompt plus expected output — see Context Windows
  • Structured output or JSON schema support — see Structured Outputs
  • Tool / function calling — see Function Calling and Tool Calling
  • Vision or other modalities
  • Reasoning depth the product has actually measured
  • Language coverage

A model that cannot satisfy a required capability is not a cheaper option. It is not in the candidate set.

Quality requirement

Not every request has the same cost of being wrong. Tag extraction for an internal search index is not the same as a customer-facing legal summary. Quality is not “use the largest model.” It is a product requirement: what failure looks like, and whether a cheaper model meets the bar on a representative eval set.

Latency

Interactive paths have a deadline the user can feel. Batch paths can wait. A router that ignores latency will send UI traffic to a slow reasoning model because it scores higher on a quality benchmark. Pair this signal with Latency Optimization: the routing decision itself should be cheap compared with generation.

Cost

Budget-aware routing and cost ceilings are legitimate signals. They are not the only signals. Cost Optimization is mostly smaller prompts, caching, and sending work to an adequate model — not a gateway tax. A cost ceiling that can override a required capability or a tenant restriction is a policy bug.

Availability

Provider and model health, capacity, and rate limits are operational signals. Clearly unhealthy or unavailable candidates can be excluded from the eligible set. Capacity and load can also influence ranking among those that remain. A health check does not guarantee the provider will stay healthy for the actual request. Availability should not silently rewrite a quality-sensitive route without recording that a fallback ran.

Tenant or policy constraints

Allowed providers, regions, data-handling requirements, and contractual restrictions are hard constraints. They are closer to AI Security than to model preference. A router that can pick a disallowed provider because it is cheaper has bypassed the security boundary.

Request characteristics

Token size, language, modality, and known complexity (document length, number of tools, schema strictness) are features of this request. They often determine context-window eligibility and whether a “small” model can physically accept the payload.

Evaluation results

Offline and periodic eval can attach quality, latency, cost, and failure-rate priors to a route. That is how the system answers “is Model A actually better than Model B for this workload?” without relying on launch-week anecdotes. See Routing with Evaluation.

Engineering Insight

Unused signals are not a defect. A router with three reliable inputs and an explainable policy outperforms a router that consumes twenty noisy features.

Routing Strategies

Strategies range from a pin to a multi-stage policy. Later strategies are not automatically better.

A. Static / pinned routing

A known task always uses a known model. support_classifier → Model A. Version the pin. Log the model ID. This is the correct design when the workload is uniform.

B. Rule-based routing

Explicit policies determine the model: if/else, tables, or config. Rules can combine task, payload size, and tenant. They remain inspectable. Most production routers should start and often stay here.

C. Capability-based routing

Choose from models that satisfy required capabilities, then apply preferences. The catalog needs an honest capability matrix. If the matrix is wrong, the router will be confidently wrong.

D. Cost/latency-aware routing

Among acceptable models, prefer lower cost or lower latency. “Acceptable” is the important word. This strategy is a ranker on an already-filtered set, not a replacement for capability and policy filters.

E. Quality-aware routing

Use measured quality — golden-set scores, human review, or online outcome signals — to influence selection. Quality-aware routing without evaluation is just preference dressed as evidence.

F. Dynamic / classifier-based routing

A separate classifier or lightweight model estimates which route is appropriate, typically when the task class is not known up front or when complexity varies inside one product surface. This can be useful. It is not inherently superior. The classifier is another model to evaluate, version, and fail. If a request already knows it is a summarization job, do not classify it.

A learned router can go further: it uses historical preference or evaluation data to predict which model is likely to offer the best quality/cost trade-off for this request. RouteLLM is one example of that approach. It remains optional, and it still has to respect hard constraints.

G. Multi-stage routing

First filter by hard constraints, then rank remaining candidates on soft preferences. This is a natural production default once more than one signal exists. It is still deterministic if the filters and the scoring function are deterministic.

Strategy When it is enough Extra failure mode
Pinned One workload, one adequate model Hidden when a second workload appears
Rule-based Known tasks and requirements Rule sprawl if every exception becomes a clause
Capability-based Mixed modalities or APIs Stale capability metadata
Cost/latency-aware Eligible set already quality-safe Optimizing the wrong objective
Quality-aware You can measure the workload Overfitting a tiny eval set
Classifier-based Task is genuinely unknown Router errors send work to the wrong tier
Multi-stage Hard constraints plus preferences coexist Complexity if stages are not named

Important

Deterministic routing is often preferable when requirements are known and predictable. Dynamic or LLM-based routing is a tool for residual uncertainty, not a replacement for a policy you can already write down.

Hard Constraints vs Soft Preferences

Production routing policies almost always contain two different kinds of rule. Mixing them is a common source of security and quality incidents.

Hard constraints eliminate candidates. A violation means “this model must not serve this request,” regardless of cost or speed.

Examples:

  • Provider not allowed for this tenant
  • Required capability missing (vision, tools, structured output)
  • Region or data-residency restriction
  • Context window insufficient for the payload
  • Tenant policy or contractual restriction

Soft preferences rank the candidates that remain.

Examples:

  • Lower cost
  • Lower latency
  • Preferred provider
  • Higher measured quality
  • Better historical performance on this task

A router should eliminate candidates that violate hard constraints before optimizing soft preferences. If the remaining set is empty, fail closed or use an explicit, policy-approved fallback — do not “relax” a tenant restriction to save the request.

Simple example:

  1. Request: summarization, tenant acme, ~8k input tokens, interactive SLO.
  2. Hard filter: drop providers acme cannot use; drop models with context window below the prompt; drop models without the summarization capability you actually require.
  3. Remaining: Model B and Model C.
  4. Soft rank: Model B meets the latency SLO at lower cost; Model C scores slightly higher on the summarization eval.
  5. Policy says latency SLO beats marginal eval gain for this product surface → select Model B.
  6. If Model B is unavailable, fallback may try Model C if it still satisfies the hard filters.

The same request must not select a disallowed provider because it is 10% cheaper. That is not ranking. That is a constraint failure.

Step-by-Step Flow

A typical production decision looks like this:

  1. Receive the request with tenant identity, payload, remaining deadline, and any explicit task class.
  2. Identify the task from a trusted field, workflow step, or — only if needed — a classifier.
  3. Apply tenant and provider restrictions so disallowed backends never enter ranking.
  4. Filter by required capabilities and context-window fit.
  5. Remove clearly unhealthy or unavailable candidates using health signals and current rate-limit state.
  6. Rank remaining candidates by the organization’s preference order (quality, latency, cost, preferred provider, and capacity where it is known).
  7. Select a model and record the reason plus policy version.
  8. Execute the request with a timeout inside the remaining deadline.
  9. Observe the result — success, error class, tokens, latency — and, if needed, run a bounded fallback.

If step 5 or 6 empties the set, the system should fail with an explicit “no eligible model” outcome rather than widening hard constraints.

Production Control-Flow Example

The following is illustrative pseudocode, not production-ready code and not a vendor SDK. It shows the control flow: filter, then rank, then execute with a separate fallback path.

# Illustrative pseudocode — not a production SDK and not a real vendor API.

def identify_task(request):
    # Routing-relevant task identity comes from trusted application context,
    # not arbitrary user input.
    return request.trusted_task


def eligible_models(request, catalog, policy, health):
    task = identify_task(request)
    candidates = catalog.for_task(task)

    allowed = []
    for model in candidates:
        if not policy.allows(request.tenant, model):
            continue
        if not model.satisfies(request.required_capabilities):
            continue
        # Full context need: input + expected output + tool/schema overhead.
        if model.context_window < request.required_context_tokens:
            continue
        if not health.is_available(model):
            continue
        allowed.append(model)
    return allowed


def route(request, catalog, policy, health):
    candidates = eligible_models(request, catalog, policy, health)
    if not candidates:
        raise NoEligibleModel("no model satisfied hard constraints")

    ranked = sorted(
        candidates,
        key=lambda model: policy.preference_key(model, request, health),
    )
    selected = ranked[0]
    return RouteDecision(
        model=selected,
        reason=policy.explain(selected, request),
        remaining=ranked[1:],
    )


def execute_with_fallback(request, decision, caller_deadline):
    attempts = [decision.model, *decision.remaining[:1]]  # bound the chain
    last_error = None
    for model in attempts:
        remaining = caller_deadline.remaining()
        if remaining <= 0:
            break
        try:
            return call_model(model, request, timeout=remaining)
        except TransientProviderError as err:
            last_error = err
            # Same-backend retry (not shown): honor Retry-After only if the
            # remaining deadline still allows that wait. Moving to the next
            # eligible backend is a separate decision and must not inherit
            # this provider's Retry-After delay.
            if not should_fallback(err):
                raise
            continue
    raise RouteExhausted(last_error)

What the sketch is trying to make obvious:

  • Task identity comes from trusted application context, not from arbitrary user input or an extra LLM by default.
  • Tenant policy and capability checks happen before ranking.
  • Health should generally gate eligibility; capacity and operational signals can also influence ranking among eligible candidates. A health check does not guarantee the provider will remain healthy for the actual request.
  • The reason is part of the decision, not a log line added later.
  • Fallback is a short, explicit list under the remaining deadline — not an unbounded retry loop. Same-backend retries may honor Retry-After when the deadline allows; fallback to another backend does not inherit that delay.

Before using a fallback candidate, production implementations should re-check dynamic eligibility signals such as current availability and rate-limit state; static policy and capability constraints remain part of the original decision.

Wire this behind whatever adapter you already use to call providers. If that adapter is a shared boundary, it is an AI Gateway. The router does not need to own credentials, HTTP retries, or token accounting.

Advanced Routing: Quality, Cost, and Latency

Once more than one eligible model exists, routing becomes a policy / optimization problem rather than a lookup table.

A router may be trying to improve several dimensions at once:

Objective Typical pressure
Quality Prefer models that score higher on the task
Latency Prefer models that meet the interactive SLO
Cost Prefer lower unit cost within the envelope
Availability Prefer backends that can accept the request now

These objectives conflict.

  • Highest quality often costs more and may be slower.
  • Lowest cost may increase latency or miss a quality bar.
  • Lowest latency shrinks the candidate set (some models cannot meet the SLO).
  • Provider availability can change the optimal route minute to minute.

There is no universal scoring function. An interactive support classifier may rank latency, cost, quality. A high-risk summary may rank quality, policy, latency, cost. The same catalog can serve both products if the policy object is per product or per tenant, not a global “best model” score.

Keep the math modest. A lexicographic order (satisfy SLO, then minimize cost) or a short weighted score on an already-filtered set is enough for most teams. Academic multi-objective optimization is optional. Explainability is not: if operators cannot say why Model B won, you cannot debug a regression.

Routing with Evaluation

Routing policies should be informed by evaluation, not by intuition alone.

A model that “feels smarter” is not a route. Compare candidates on representative workloads for each task class:

  • Quality by task (correctness, format, faithfulness where retrieval is involved)
  • Latency (p50/p95, not only mean)
  • Cost per successful request
  • Failure rates (timeouts, refusals, schema violations, provider errors)

Evaluation answers a question routing cannot answer by construction: is Model A actually better for this workload than Model B? Public leaderboards and benchmarks shortlist models. Product eval decides whether a route is safe to ship. See also LLM Evaluation for generation-level measurement and Agent Evaluation when the “model” sits inside a tool loop — a cheaper model that calls tools well may outperform a larger model that does not.

Practical consequences:

  • Pair every new route with a slice of the golden set for that task.
  • Re-evaluate when providers change versions, prices, or behavior.
  • Do not promote a cheap model on cost dashboards unless quality on the hard slice is still acceptable.
  • Traffic splitting and canaries are routing operations; they need the same eval hooks as a full cutover.

Warning

A routing change is a behavior change. Treat it like a deploy: version the policy, canary it, and keep a rollback to the previous pin.

Routing and Observability

Routing without visibility becomes folklore. You do not need a separate observability architecture on this page — see Observability for traces, metrics, and logs as a system. You do need the route to be a first-class field on the request span.

Minimum useful record:

Field Why it exists
Selected model ID What actually ran
Provider Where it ran
Routing reason / policy ID Why this candidate won
Latency Decision time versus generation time
Token usage and cost Attribution per tenant and route
Error class Timeout, 429, 5xx, capability miss, policy deny
Retries / fallbacks Whether the preferred path held
Outcome / eval signals Online quality hooks where you have them

If you cannot answer “why was this model selected for this tenant last Thursday?”, you do not yet have a routing policy. You have a default that drifted.

Prompts, completions, tool arguments, and retrieved content should not be captured automatically merely because routing is observable; they may contain sensitive data and should follow the application’s telemetry redaction and retention policy.

Keep the routing decision itself cheap and synchronous enough to trace. A router that calls a slow LLM to choose a fast LLM has already spent the latency budget it was created to save.

Design Decisions

Question Simpler choice More advanced choice When to prefer the advanced choice
One model or many? Pin one ID Catalog of task-specific models Workloads differ in capability, quality, or cost
Static rules or dynamic routing? Config / if-else Classifier or scored router Task or complexity is not known from the application
Single provider or multi-provider? One vendor SDK Provider + model catalog Availability, region, or capability gaps are real
Cost vs quality optimization? Meet a quality bar, then cut cost Multi-objective ranker You have eval coverage and conflicting SLOs
Where does routing live? Application-local function Shared policy on the AI gateway Several services would otherwise copy the same table
Fallback depth? Fail the request One named alternate Availability matters more than identical answers
How is the decision explained? Log model ID Policy version + reason code You operate more than one route in production

More complexity is not automatically better. Each extra signal is another way to be wrong and another field to keep fresh.

Common routing patterns

These are named ways to fill the table above. Use them when the situation matches, not as a checklist.

Pattern What it does When it makes sense
Cheap-first Prefer the lowest-cost eligible model Quality bar is met by the cheap model on eval
Quality-first Prefer the best measured model, ignore extra unit cost Cost of errors dominates token cost
Latency-sensitive Drop candidates that miss the SLO before ranking Interactive UI or strict p95
Capability routing Filter by vision, tools, context, schema, language Mixed modalities or APIs
Tenant-aware Allow-lists and data-handling rules first Multi-tenant or regulated products
Regional / provider routing Choose a backend that may serve this region Residency, sovereignty, or regional capacity
Fallback routing Alternate eligible model after a classified failure Provider blips without relaxing hard constraints
Traffic splitting Send a percentage to a candidate for eval Comparing models with live traffic
Gradual migration Canary from Model A to Model B by task or tenant Replacing a pin without a big-bang cutover

Cheap-first without a quality gate is how hard cases rot. Quality-first without a cost envelope is how a classifier becomes a frontier-model bill. Traffic splitting without assignment in traces makes the experiment unreadable.

Comparisons

Model routing vs model selection

Selection is catalog design: which models are offered, at which versions, for which tasks. Routing is the per-request use of that catalog. You can select models quarterly and still route every request. Confusing the two leads to “we picked Claude for the company” as if that were a request policy.

Model routing vs fallback

Routing: “Use Model B for summarization because it meets the task requirements.”

Fallback: “Model B failed or became unavailable, so try Model C.”

Fallback can be implemented next to routing operationally. It is still a different question. Fallback is triggered by timeout, provider errors, rate limits, or temporary unavailability, and it must fit the remaining request deadline. It is not an unlimited retry chain. If the preferred backend returns Retry-After, a same-backend retry should honor that delay only when the remaining deadline allows it. Switching to another eligible backend is a separate decision and should not blindly inherit the original provider’s Retry-After. A cheaper or different model is a behavior change, not a hot spare with the same answers. Pair fallbacks with evaluation if quality matters — the same warning as in the AI Gateway guide.

Diagram: Preferred route versus fallback

stateDiagram-v2
    [*] --> Route
    Route --> Execute: Selected model
    Execute --> Done: Success
    Execute --> Fallback: Timeout, rate limit, or unavailable
    Fallback --> Execute: Alternate eligible model
    Fallback --> Fail: No remaining budget or candidates
    Done --> [*]
    Fail --> [*]

Fallback starts after the preferred route is known. It does not replace the initial decision, and it must stop.

Model routing vs load balancing

Load balancing spreads traffic across equivalent replicas (same logical model, multiple instances or keys). Routing chooses among non-equivalent models based on request requirements. After a model is selected, load balancing may still distribute traffic across equivalent replicas of that model. Using a load balancer as a capability router will send vision requests to text replicas whenever those replicas are idle.

Model routing vs orchestration

Orchestration sequences steps: retrieve, call tools, generate, validate. Routing chooses the model for a generation step. An agent framework can contain a router; it is not a router by itself. See AI Agents and AI System Architecture.

Provider routing vs model routing

Model routing asks: which model should handle the request?

Provider routing asks: which provider should serve it?

They combine. A logical requirement such as “interactive summarization, 32k context, tenant-allowed” may have several provider-specific model IDs. Separating the logical capability from the provider implementation is useful: you can fail over providers without changing the product’s idea of the task, and you can change a model ID without pretending the task changed.

A gateway often performs provider adapters; the router decides which adapter and which ID. See AI Gateway for that boundary. Do not duplicate credential, authentication, rate-limit, or adapter design here.

Model routing and the AI Gateway

Conceptual boundary:

Application → AI Gateway → routing / policy → provider / model

The gateway can provide the provider-access and policy enforcement point. Routing determines which eligible model or provider should handle the request. Either can exist without the other: a single-provider app can route between tiers in application code with no gateway; a gateway can exist solely for credentials, timeouts, and traces with one model behind it.

Details of credentials, caller authentication, rate limiting, observability pipelines, and provider adapters belong in AI Gateway. This guide only needs the split: boundary versus decision.

Routing Policy and Precedence

Conflicting requirements need an explicit order. One common — not universal — precedence is:

  1. Security / tenant restrictions
  2. Required capability
  3. Context-window requirements
  4. Reliability / availability
  5. Quality target
  6. Latency target
  7. Cost preference

Organizations define their own order. A latency-critical classifier may place SLO above measured quality. A regulated workload may place region above availability (better to fail than to leave the region). The key concept is that routing should be explainable: the system should be able to answer “why was this model selected?”

Store precedence in the policy, not in tribal knowledge. When two teams disagree about cost versus quality, that is a product decision to encode, not a reason to add another classifier.

Common Mistakes

  1. Routing based only on model price. The cheap model that cannot see the image, fill the schema, or meet the quality bar is not a saving.

  2. Assuming the largest model is always best. Extra reasoning depth does not help a two-class intent label. It does add latency, spend, and rate-limit risk.

  3. Using an LLM router when deterministic rules are sufficient. If the application already knows the task, a map is simpler, faster, and easier to evaluate.

  4. Allowing routing to bypass tenant or security policy. Soft preferences must not resurrect a candidate the hard filter removed.

  5. Treating fallback as unlimited retries. Completions are often billable per attempt. Cap the chain, stay inside the caller deadline, and honor Retry-After only when retrying the same backend if that wait still fits.

  6. Ignoring context-window limits. Sending an over-long prompt to a small model is not routing. It is a guaranteed failure that should have been filtered.

  7. Routing without measuring outcomes. A policy that is never eval’d will drift as providers change versions underneath stable IDs — or as IDs are swapped without a golden set.

  8. Creating overly complicated routing rules. Every special case is a branch that will not be tested. Prefer a small table plus an explicit exception list.

  9. Making routing decisions impossible to explain. If the only artifact is “the router chose B,” you cannot tell a policy bug from a provider outage.

  10. Changing routing policies without regression evaluation. A canary that does not include the hard slice of the task will look green.

Where It Breaks Down

Poor task classification. If the task label is wrong, every downstream filter is applied to the wrong catalog slice. Trusted application labels beat inferred labels.

Insufficient evaluation data. Quality-aware and cheap-first policies both need a representative set. A dozen happy-path prompts will not detect the long-document failure.

Rapidly changing provider and model behavior. Versionless aliases move under you. Pin IDs, re-eval on change, and treat silent alias updates as deploys.

Correlated provider failures. Multi-provider fallback does not help if both providers share a region, a GPU shortage, or the same upstream dependency. Diversity on paper is not diversity in the incident.

Hidden quality regressions. A cheaper route can pass format checks and fail usefulness. Online feedback loops that only measure thumbs-up will hide task-specific damage.

Routing complexity. A policy no one can simulate in staging will not be debugged in production. Complexity is a reliability hazard.

Stale routing policies. Capability matrices, prices, and context windows go stale. A quarterly catalog review is part of running the router.

Feedback loops. If online quality signals train the classifier that chooses the route, a biased sample can lock traffic onto a worse model. Keep a held-out eval path.

Unpredictable workloads. If every request is a new shape, static routes underfit and classifiers overfit. You may want a stronger default model and fewer routes, not more machinery.

Routing is only as good as the signals and measurements behind it. When those are weak, pin fewer models.

When NOT to Use Model Routing

A single-model application is often the correct architecture.

Skip a multi-model router when:

  • One model already meets quality, latency, and cost for the actual workload
  • You cannot name how tasks differ in a way a second model would improve
  • You do not yet measure quality, so a second route would be an untested hypothesis
  • The “router” would exist only because an architecture diagram included one
  • You would need an LLM to classify work that the application already labels

Use routing when several of these are true:

  • Multiple models genuinely provide different value (capability, quality, or cost)
  • Workloads differ materially by task, modality, or risk
  • Cost, latency, or availability tradeoffs are real and measurable
  • Provider diversity is an operational requirement you can test
  • Migration or experimentation needs controlled selection (canary, split, rollback)

Do not introduce routing merely because it sounds architecturally sophisticated. An unmeasured mesh of models is harder to operate than a pinned ID with traces.

Decision tree: do you need more than a pin?

flowchart TD
    Start[More than one model or provider?] -->|No| Pin[Pin one model ID]
    Start -->|Yes| Diff{Do tasks differ in capability, quality, latency, or cost?}
    Diff -->|No| Avail{Need availability failover?}
    Avail -->|No| Pin
    Avail -->|Yes| Fallback[Named fallback]
    Diff -->|Yes| Known{Are requirements known?}
    Known -->|Yes| Rules[Deterministic routing]
    Known -->|No| Measure[Evaluate then add rules]
    Measure --> Still{Still unpredictable?}
    Still -->|No| Rules
    Still -->|Yes| Class[Optional classifier]

Start from a pin. Add rules when jobs differ. Add a classifier only when the task is not already known.

Running in Production

Best Practice

Keep decision logic bounded and deterministic where possible. Filter hard constraints first. Cap fallbacks. Version the policy. Emit the reason. Re-evaluate routes on a schedule, not only after incidents.

Dimension Guidance
Decision logic Bounded, testable functions or tables — not an unbounded agent loop
Determinism Prefer rules you can replay in staging with the same inputs
Deadlines Routing + generation + fallback must fit the caller’s remaining budget
Health Use availability signals; do not wait forever to discover a 429
Fallback Named, short, constraint-preserving; honor Retry-After on same-backend retries when the remaining deadline allows — do not inherit that delay onto another backend
Observability Model, provider, policy version, reason, tokens, errors, fallback flag
Evaluation Golden set per task class; block policy deploys on regressions
Versioning Policy IDs like prompt and index versions
Rollout Canary by task or tenant; keep an instant rollback to the previous pin
Security Tenant restrictions are not optional ranker features

Production checklist

  • Every request records model_id, provider, and policy version
  • Hard constraints (tenant, region, capability, context window) run before ranking
  • Soft preferences cannot resurrect a disallowed candidate
  • Timeouts on every provider call; fallback stays inside the remaining deadline
  • Fallback chain is capped and logged; same-backend retries honor Retry-After when the remaining deadline allows it; fallback to another backend does not inherit that delay
  • Each route has a representative eval slice; policy changes go through it
  • Health / rate-limit signals can mark a model temporarily ineligible
  • Operators can answer “why this model?” from traces
  • A single-model pin remains a supported configuration
  • Alias-free model IDs, or aliases treated as deploys when they move

Interview Questions

  1. What is model routing?
    The decision process that chooses which eligible model or provider should handle a request. It can be a static map. It does not require an LLM router.

  2. How is routing different from fallback?
    Routing is the preferred choice given current signals. Fallback is an alternative path when that choice cannot be used or fails. Fallback must be bounded.

  3. When would deterministic routing be better than an LLM router?
    When the application already knows the task and requirements. Rules are faster, cheaper, testable, and explainable. Use a classifier when the task is genuinely unknown.

  4. How would you route between models with different cost, latency, and quality?
    Filter by hard constraints first. Rank the remainder with an explicit precedence (for example: meet quality bar, meet SLO, then minimize cost). Measure all three on a golden set.

  5. How would you enforce tenant restrictions?
    As a hard filter before ranking, using authenticated tenant identity. Never as a soft preference the ranker can override.

  6. How would you know whether a routing policy is actually working?
    Per-route eval (quality, latency, cost, errors), traces that include the reason, and regression gates on policy changes. Dashboards that only show spend are not sufficient.

  7. Where should routing live relative to an AI gateway?
    The gateway is the provider-access boundary. Routing is a policy that may run there or in the application. Either can exist without the other. See AI Gateway.

  8. How would you safely change routing rules in production?
    Version the policy, canary by task or tenant, compare against the previous pin on the golden set, and keep an instant rollback. Treat provider alias changes the same way.

This guide is the decision process. Production AI Stack is the capability map. AI Gateway is the provider-access boundary. AI System Architecture is the platform blueprint.

Architecture:

Foundations:

Operations:

Agents (routing is not orchestration):

Rankings: Best AI APIs — hosted access options, not a mandate to route across all of them

Tools: LiteLLM — one implementation of adapters and routing tables, not a required architecture

Learning path: Become an AI Engineer

Diagram: Recommended reading around this guide

flowchart LR
    Stack[Production AI Stack]
    GW[AI Gateway]
    Route[Model Routing]
    Eval[Evaluation]
    Stack --> GW
    Stack --> Route
    GW --> Route
    Route --> Eval

Read the stack map for composition, the gateway guide for the access boundary, and evaluation for whether a route is actually good.

Learning Path

Prerequisites: Large Language Models

Next topics: AI Gateway · Production AI Stack · AI Copilot Architecture · Evaluation · Cost Optimization · Observability

Estimated time: 40 min · Difficulty: Advanced

FAQs

Is model routing the same as model selection?

No. Selection is deciding which models belong in the catalog (design-time or eval-time). Routing is deciding which eligible catalog entry handles this request. You need both; they are not the same job.

Does model routing require an LLM?

No. A task-to-model map is routing. An LLM or classifier router is optional when the task is not already known. It is not the definition.

Is model routing the same as fallback?

No. Routing chooses the preferred eligible model. Fallback is used when that path cannot be used or fails. A system may implement both together; the questions remain different.

Should routing happen inside the AI gateway?

It can. The gateway is a natural enforcement point when several services share policy. Application-local routing is correct when one service owns the workloads. See AI Gateway.

How do you choose between cost and quality?

Encode a bar, not a vibe. Filter to models that meet the quality requirement on eval, then optimize cost (or latency) among those that remain. If no model meets the bar at the allowed cost, that is a product decision — not a ranker tie-break.

When should a system use multiple providers?

When a second provider supplies a capability, region, or availability property the first cannot, and you can test failover as a behavior change. Multiple providers are not required for routing between tiers of one provider.

How often should routing policies be reevaluated?

Whenever models, prices, context windows, or quality bars change — and on a regular cadence even if nothing was announced. Alias-based model IDs can move without a deploy.

Is load balancing a substitute for routing?

No. Load balancing assumes equivalent backends. Routing assumes they are not equivalent. You may load-balance replicas of the model you just selected.

Can routing live in the application instead of a platform service?

Yes. A function that returns a model ID from a table is a router. Extract a shared policy when several codebases copy the same rules.

What if no candidate survives hard constraints?

Fail closed, or use a pre-approved degraded path that still satisfies policy (for example a refusal message). Do not relax tenant or residency rules to obtain a completion.

Should every request go through a complexity classifier?

Only if complexity is not already known and eval shows the classifier improves outcomes enough to pay for its errors and latency. Many systems already know the job type.

References

Further Reading

Key Takeaways

  • Model routing is the per-request decision that chooses an eligible model or provider — a static task map already counts.
  • It exists because models differ in capability, quality, latency, cost, availability, and policy; neither “always strongest” nor “always cheapest” is a production strategy.
  • Hard constraints filter; soft preferences rank. Tenant and capability rules are not optional score terms.
  • Routing ≠ fallback ≠ gateway ≠ evaluation. The gateway is a boundary; fallback is recovery; eval tells you whether the route works.
  • Deterministic routing is often the right design; an LLM router is for residual uncertainty, not prestige.
  • Observe model, provider, reason, and policy version, and re-evaluate routes when models or workloads change.
  • A single pinned model is a valid architecture. Add routing when mixed workloads actually require it.

Next Topics

Learning Path

Continue Learning

Related Guides

Related Tools

ToolCategoryPurposeWebsiteBest For
LiteLLM
Open SourceAPI
infrastructureUnified API gateway for 100+ LLM providers with routing and fallbacks.litellm.aiMulti-provider routing

Related Rankings