Architecture

Production AI Stack Guide

What a production AI application needs beyond a model and a prompt — models, gateway, routing, retrieval, agents, evaluation, observability, and runtime as composable capabilities.

40 min readAdvancedLast reviewed: 2 September 2026

Quick Summary

The production AI stack is the set of application capabilities — model access, retrieval, agents, evaluation, observability, and runtime — that turn a model call into a system you can operate.

One Analogy

A model is an engine; the production AI stack is the rest of the vehicle — steering, brakes, instruments, and the road the engine never sees.

Engineering Rule

Treat the stack as capabilities you compose, not a mandatory platform. Add a layer when a failure mode appears, not when a vendor diagram includes it.

Production AI Stack in One Sentence

Production AI Stack =

  • Models
  • AI Gateway
  • Model Routing
  • RAG / Retrieval
  • Agents
  • Evaluation
  • Observability
  • Production Runtime

TL;DR

  • A production AI application is software that calls models, not a model with a UI. Prompts, weights, and a chat box do not give you retries, authorization, grounding, evaluation, or an incident trail.

  • The stack is a set of capabilities, not a shopping list. Simple generation may need only a model, an application, and basic tracing. Retrieval, gateways, agents, and heavy eval are responses to specific problems.

  • Layers are architectural roles, not required services. An AI gateway can be a library, a proxy, or a few functions in your API. Model routing can live in application code. Agents are an orchestration pattern, not a separate product you must buy.

  • Compose along failure modes. Private or changing knowledge → retrieval. Multiple models or providers → gateway and routing. Open-ended tool use → agents with bounds. High-risk or fast-changing behavior → stronger evaluation and observability.

  • Read this as a map, then go deep elsewhere. Layered platform design lives in AI System Architecture. The provider-access boundary is AI Gateway. RAG, agents, eval, and observability each have dedicated guides — this document explains how those pieces fit and when they are worth the cost.

Architecture Snapshot

Complexity

★★★★☆

Audience

AI Engineers, Tech Leads, Architects

Difficulty

Advanced

Typical Deployment

Production application

Typical Latency

Workload-dependent

Scalability

Add layers as failure modes appear

Availability Target

Match the surrounding product

Read Time

~40 min

Last Updated

September 2, 2026

Recommended Stack

  • Application runtime (API + jobs)
  • Model access (SDK or gateway)
  • Retrieval / search when grounding is required
  • Bounded agent/tool loop when paths are open-ended
  • Golden-set eval + request traces

Why This Matters

A prototype is usually one call: serialize a prompt, hit a provider, return text. That shape hides most of the work that shows up the week after launch.

Production traffic is concurrent. Knowledge goes stale. Providers return 429s and 503s. The same question must sometimes hit a cheap model and sometimes a slower one. Answers that sounded fluent in a demo fail a compliance review. A customer reports a wrong answer from Tuesday and nobody can reconstruct the prompt, the retrieved chunks, or the tool calls.

Those are application problems. The model does not own retries, tenant isolation, citation checks, or cost attribution. If you do not name the capabilities around the model, they accumulate as glue in request handlers — untested, unobserved, and expensive to change.

This guide answers a practical question: what does a production AI system actually need beyond a model and a prompt? It is a composition map. The companion AI System Architecture guide is the layered platform blueprint (ingestion, orchestration, contracts, deployment). Use this page to decide which capabilities you need; use that page when you are ready to draw service boundaries.

The Problem a Production Stack Solves

Three failure modes appear when teams ship the prototype unchanged.

The model is treated as the system. Business logic, retrieval, tool permissions, and output policy sit in prompt strings. You cannot test them independently, you cannot scale them independently, and you cannot tell which part broke.

Capabilities are copied from a vendor diagram. A “full stack” appears: vector database, agent framework, gateway, eval SaaS, feature store. Most of it is idle. Complexity shows up as latency, cost, and incident surface before it shows up as quality.

There is no place to put production concerns. Retries live in one helper, cost caps in another, tracing in a third. When a request is slow, you cannot say whether retrieval, the provider, or a tool hung.

A production stack names the roles — model access, retrieval, orchestration, evaluation, observability, runtime — so you can adopt them in order, keep them thin, and avoid pretending every application is an enterprise platform.

Without a stack map With a stack map
Prompt file + SDK call is “the architecture” Each capability has a job and a reason to exist
Every new feature adds another vendor You add a layer when a failure mode appears
Incidents start at “the LLM is wrong” You can attribute retrieval, routing, tools, or runtime
Eval is a spreadsheet after launch Quality gates sit next to deploys

How We Got Here

Production AI did not arrive as a new category of infrastructure. It is ordinary application architecture with a probabilistic component in the request path.

Diagram: A simplified evolution of production AI architectures

timeline
    title Simplified architectural evolution
    2020-2022 : Provider APIs
              : Prompt in the handler
    2023 : Chat UIs over one model
         : RAG demos and vector DBs
    2024 : Gateways and eval harnesses
         : Tool-calling agents
    2025-2026 : Compose only what you need
              : Traces, golden sets, bounds

A simplified teaching sequence — not a literal industry chronology.

Era What shipped What broke in production
Single API call Chat wrapper over one provider No grounding, no failover
RAG prototypes Embed, retrieve, stuff context Authz, freshness, retrieval quality
Framework chains Fast composition Opaque control flow, buried cost
Agent demos Tools in a loop Unbounded steps, irreversible actions
Operable systems Routing, traces, golden sets, bounds Requires deliberate composition

The useful lesson is not “collect every box.” It is that models are interchangeable; the application around them is not. Gateway products, vector databases, and agent runtimes are implementations of roles. Your job is to decide which roles you have earned.

What Is the Production AI Stack?

The production AI stack is the set of capabilities an application uses to call models safely, ground them when needed, take actions when needed, and remain operable after deploy.

It is not a reference cluster you must stand up. It is not a ranking of products. It is a vocabulary for the work around the model:

  • Models — what generates or classifies
  • AI gateway — a common way to reach providers
  • Model routing — which model handles this request
  • RAG / retrieval — how external knowledge enters the prompt
  • Agents — how the system chooses tools and next steps
  • Evaluation — how you know quality did not regress
  • Observability — how you reconstruct a request
  • Production runtime — the application that hosts all of the above

Two distinctions matter.

Capability versus service. Retrieval can be a PostgreSQL extension in the same process as the API. An “AI gateway” can be a 200-line proxy or a provider SDK with retries. Agents can be a for loop with a tool table. Do not start by drawing microservices.

This stack versus AI System Architecture. This guide is what to compose. That guide is how to layer a full platform — ingestion versus query, orchestration as control plane, contracts, and operational boundaries. Specialized production patterns (Enterprise RAG, GraphRAG, Knowledge Graph + LLM) sit on top of both.

The Production AI Stack

At a glance, a production AI application is a runtime that may use some or all of the following. Nothing below the application is mandatory.

Layer Question it answers You need it when
Models What can generate, classify, embed, or extract? Always — this is the non-negotiable core
AI gateway How do we talk to providers uniformly? Multiple providers, shared credentials, or central policy
Model routing Which model (or tier) for this request? Quality, cost, latency, or availability differ by task
RAG / retrieval What evidence is not in the weights? Private, current, or citable knowledge
Agents What should we do next, including side effects? The control flow is not a fixed pipeline
Evaluation Did this change make answers worse? You ship prompts, retrieval, or tools more than once
Observability What happened on request X? You have users, SLOs, or a cost bill
Production runtime Where does this run, persist, and fail? Always — this is still an application

Gateway and routing overlap: routing is often a policy the gateway executes. They are listed separately because a single-provider app can route between model tiers without a gateway, and a gateway can exist solely for credentials and retries with one model behind it.

Architecture

Diagram: Production AI stack

flowchart TB
    Runtime[Production Runtime]
    Runtime --> Models
    Runtime --> Gateway[AI Gateway]
    Gateway --> Routing[Model Routing]
    Routing --> Models
    Runtime --> RAG[RAG / Retrieval]
    RAG --> Search[Vector / Search]
    Runtime --> Agents
    Agents --> Tools[Tools / Memory]
    Runtime --> Eval[Evaluation]
    Runtime --> Obs[Observability]

The runtime is the application. Other boxes are capabilities it may call — not a required topology.

The diagram is a map of roles. Edges are “may use,” not “must deploy.” A simple generator is Runtime → Models, with observability as logs and traces on that path. A RAG assistant adds retrieval. An agent assistant adds a bounded loop over tools. Evaluation is usually offline-plus-CI, not a box on the hot path — except when you sample production traces into an eval set.

Engineering Insight

If you cannot explain a layer as a failure mode you have already hit (or will hit this quarter), it is decoration. Architecture diagrams are not procurement checklists.

1. Models

The model layer is the foundation: one or more models that generate text, classify intent, embed queries, extract structure, or score candidates.

Production applications rarely use a single model for everything. Embedding models are not chat models. A cheap classifier that routes “reset password” away from a frontier model is often a better architecture than a smarter prompt. Vision, speech, and rerankers are additional model roles, not features of “the LLM.”

Selection is an engineering decision, not a leaderboard decision. You care about:

  • Capability — reasoning, tool calling, long context, structured output, multilingual
  • Cost — tokens in and out, embedding volume, rerank calls
  • Latency — time to first token and time to last token under your concurrency
  • Context — window size versus the documents and history you actually send
  • Control — data handling, region, rate limits, deprecation policy

Large Language Models covers how these systems work. Family guides (GPT, Claude, Gemini, Llama, Mistral, DeepSeek) are for model-specific behavior. Tokens and Context Windows are the budgets the rest of the stack must respect. Cost Optimization and Latency Optimization are how this layer is operated.

Hosted access is compared independently in Best AI APIs. That ranking is about APIs, not about whether you need a gateway.

Decision Trade-off

One strong model simplifies operations and eval. Multiple models reduce unit cost and tail latency if — and only if — you can route correctly and measure quality per route. Routing without eval is a cost optimization that silently degrades answers.

You do not need a model garden on day one. You need a pinned model ID, a way to change it without rewriting handlers, and an eval set that tells you the change was not a regression.

2. AI Gateway

An AI gateway is a common interface in front of model providers. Architecturally it is an adapter plus policy: one request shape, many backends.

Typical responsibilities, when you actually need them:

  • Uniform API — chat completions, embeddings, and (sometimes) rerank behind one client
  • Credentials, quotas, and policy — centralize these when multiple services need the same provider controls
  • Retries and fallbacks — 429/503 handling, secondary provider, and, where duplicate work/cost is acceptable, hedged requests
  • Usage and cost controls — per-key budgets, max tokens, block lists
  • Telemetry — one place to emit provider latency, status, and token counts

A gateway is not mandatory. A single-provider application with one model and moderate traffic can call the provider SDK, wrap retries, and emit traces from the application. The gateway earns its keep when provider diversity, shared credentials, or central policy would otherwise be copied into every service.

Do not confuse an AI gateway with your public API gateway. The public gateway authenticates users and rate-limits clients. An AI gateway is the adapter toward model providers — it can centralize provider credentials and spend controls when those belong in one place. They can be the same process; they are different jobs.

This is not a product guide. Implementations range from a thin internal proxy to libraries such as LiteLLM. If you only need retries and a second model ID, write that in the application. If ten services all need the same provider failover, extract a gateway.

The dedicated AI Gateway guide covers when that boundary is worth introducing, what belongs there, and what does not. Related operational concerns also live in Cost Optimization, AI Security, and Observability.

3. Model Routing

Model routing is the policy that sends different requests to different models. It is often implemented inside a gateway, but it is a separate idea: the gateway is how you call models; routing is which model you call.

Applications route on:

  • Task — classify, extract, generate, embed, rerank
  • Quality — hard reasoning versus template filling
  • Latency — interactive chat versus batch
  • Cost — default cheap, escalate on low confidence
  • Availability — failover when a provider is dark
  • Context — a long document needs a long-context tier; a one-line FAQ does not

Routing can be a static map (intent → model_id), a small classifier, or a cascade (try cheap, escalate if the eval proxy is low). The policy belongs in application or gateway config, versioned and logged, so you can answer “which model served this tenant last Thursday?”

A single-model app still “routes”: every request goes to one pinned ID. Name that pin. The day you add a second model, you already have a routing table instead of a hunt through handlers.

Warning

Routing on cost without a quality gate will eventually send the hard cases to the cheap model. Pair every new route with a slice of the evaluation set.

The dedicated Model Routing guide covers how that decision is made — from a static task map to production policies involving capability, quality, latency, cost, availability, fallback, and evaluation. Treat routing as a versioned policy next to Cost Optimization and Latency Optimization, enforced where you already emit traces.

4. RAG and Retrieval

Retrieval exists because weights are the wrong store for private, current, or citable knowledge. RAG (retrieval-augmented generation) is the pattern: fetch evidence, then generate with that evidence in context.

Architecturally, RAG is a data plane the application calls before (or during) generation:

  1. Turn the user question into one or more search queries (Query Transformation when naive queries fail).
  2. Retrieve candidates from a vector database / vector search index, often fused with lexical search (Hybrid Search).
  3. Enforce authorization before retrieval and carry the resulting tenant/ACL constraints into the retrieval query (Metadata Filtering) — never rely on the prompt to enforce access.
  4. Optionally re-rank a wider candidate set down to what fits the context window.
  5. Build a prompt that cites sources; generate; check faithfulness.

Embeddings and Embedding Models are how text becomes searchable. Chunking Strategies decide what “a document” means to the index. None of this replaces the model; it changes what the model is allowed to see.

You do not need RAG for a generator that only needs parametric knowledge (tone, format, coding assistance on public APIs). You do need it when wrong answers come from missing files, stale facts, or an inability to cite.

Quality of this layer is measurable. Retrieval Evaluation scores the index; RAG Evaluation attributes failures to retrieval versus generation. The Basic RAG, Hybrid RAG, Reranking, and Retrieval Evaluation labs show those pipelines as traces — they are not a required companion to this architecture guide.

When RAG must survive tenants, ACLs, and SLOs, the production envelope is Enterprise RAG Architecture. When answers need multi-hop structure, GraphRAG / GraphRAG Architecture and Knowledge Graph + LLM are different knowledge planes, not “more vectors.”

Compare stores independently in Best Vector Databases only after you know you need a retrieval plane.

5. Agents

An agent is an orchestration pattern: the model proposes the next action, the application executes it, and the loop continues until a stop condition. A RAG pipeline is usually a fixed retrieve-then-generate path. You add agentic control when the sequence of steps is not known in advance and the system must use tools, branch, or retry.

AI Agents is the concept guide. Architecturally, agents add:

  • Tool calling / function calling — the model names a tool and arguments; your code validates, authorizes, and executes
  • Workflow versus loop — Workflows vs Agents: if the path is a DAG, use a workflow; if the model must choose the path, use a bounded loop
  • State — what is in the messages versus what is in a store
  • Memory — facts that must survive a session (and go stale)
  • Human-in-the-loop — approval gates for irreversible actions
  • Durable execution — waits, retries, and resumes measured in hours, not milliseconds
  • Protocol-shaped tools — Model Context Protocol standardizes tool I/O; it does not replace permissions or tracing
  • Reliability — Agent Evaluation scores outcomes and trajectories separately; Guardrails constrain tools and outputs

Agents are expensive: extra model calls, extra latency variance, extra ways to duplicate side effects. Cap steps, cap tokens, require idempotency on writes, and do not let the model be the authorization layer.

Agentic RAG is the hybrid: the agent may retrieve more than once, or retrieve as a tool, instead of a single retrieve-then-generate pass. Use it when one retrieval is not enough — and measure it, because extra hops are extra failure modes.

Hands-on traces: Tool Calling, Agent Evaluation, Memory. Those labs map to their own guides; this stack page does not own a Lab.

Frameworks compared in Best AI Agent Frameworks implement the loop. They are libraries behind an application interface, not a substitute for timeouts, tool allow-lists, and eval.

Best Practice

Prefer a fixed pipeline or an explicit workflow until tool choice must be dynamic. Autonomy is a cost and a reliability tax — spend it on tasks that are actually open-ended.

6. Evaluation

Traditional software tests assert deterministic outputs. Model outputs are distributions. Evaluation is the regression suite for that reality: a versioned golden set, layered metrics, and a gate that can block a deploy.

In a production stack, evaluation is not one score. It is several questions:

Question Where it lives
Did the model follow the spec? LLM Evaluation, Prompt Evaluation
Did we retrieve the right evidence? Retrieval Evaluation
Was the answer faithful and useful? RAG Evaluation
Did the agent take a sensible path? Agent Evaluation
Did a public leaderboard move? Benchmarks — shortlist only

Regression testing means the golden set is checked in CI when you change prompts, models, chunking, or tools. Production feedback means sampled traces, thumbs-down, and corrections flow into that set — otherwise you are testing yesterday’s product.

Evaluation is how you know a routing change, a new retriever, or a cheaper model did not quietly fail. Without it, every other layer is an untested hypothesis. Hallucination Detection and Guardrails are complementary: they catch some failures online; eval tells you the system got worse before users do.

You do not need a research-grade harness for a weekend prototype. You need one the moment you will change the system more than once and still claim quality.

7. Observability

Observability is request-scoped visibility: what you need to reconstruct and attribute a failure. It is not the same as evaluation. Eval answers “is quality acceptable?” Observability answers “what happened on this request?”

For an AI application, a useful trace includes:

  • The user request — identity, tenant, route, payload size
  • Model calls — model ID, prompt version, tokens in/out, finish reason
  • Latency — per hop (gateway, retrieval, generation, tools), not one blob
  • Errors — provider status, timeouts, validation failures
  • Usage — tokens and estimated cost, attributed to tenant/route
  • Retrieval — query, filters, hit count, chunk IDs (not necessarily raw text in the log store)
  • Tool calls — name, arguments hash, duration, success/failure
  • Parent/child spans — so an agent loop is a tree, not a pile of log lines

If you cannot answer “why was this slow?” or “why did this say X?” from a trace_id, you do not have production observability yet — you have stdout.

Keep instrumentation boring: OpenTelemetry-style traces plus structured logs. Product-specific UIs are optional. Do not build an “AI observability platform” as a prerequisite for shipping; emit spans from day one and add backends when you have volume.

Cost Optimization and Latency Optimization depend on these signals. AI Security depends on not putting secrets and raw PII into the trace store.

8. Production Runtime

The runtime is the application and infrastructure that hosts the stack. It is the part teams already know how to build — and then forget when the demo is a notebook.

At architecture level:

  • APIs and services — the synchronous request path (HTTP, RPC, streaming)
  • Queues and jobs — indexing, eval batches, long agent runs, webhook fan-out
  • Persistence — sessions, audit logs, documents, index versions, golden sets
  • Caching — exact and semantic caching; see Caching
  • Secrets — provider keys, tool credentials; never in prompts or client bundles
  • Scaling — stateless app replicas versus retrieval replicas versus job workers
  • Reliability — timeouts, bulkheads, backpressure, idempotency, deploy rollback

This layer is why “just call the model” fails: the model is one dependency among databases, queues, and identity. AI System Architecture goes deeper on ingestion versus query paths and orchestration as a control plane. AI Security belongs here as threat modeling, not as a system prompt.

If you already run a production web service, you already have most of this. The AI-specific work is pinning model/index/prompt versions on every request, bounding non-deterministic loops, and treating providers as unreliable backends.

How the Pieces Fit Together

The stack is compositional. Four patterns cover most products. None of them require every layer.

Simple AI application

Application → Model → Evaluation / Observability

A formatter, classifier, or writing aid with no private corpus and one provider. The runtime is an API. Observability is traces with model ID and token counts. Evaluation is a small golden set on format and a handful of quality cases. No gateway, no RAG, no agent loop.

RAG application

Application → Gateway → Model
  ↘ Retrieval → Vector / Search → Reranking

Grounded Q&A over your documents. Retrieval is on the hot path; the gateway is optional until you have more than one model or provider. Evaluation must include retrieval metrics, not only “the answer sounds good.” See RAG and Enterprise RAG Architecture.

Agent application

Application → Gateway → Model
  ↘ Agent → Tools / State / Memory

Open-ended work with side effects: tickets, queries, code, browser. The loop is bounded. Tools are permissioned in the application. Memory is an explicit store, not “whatever fit in the window.” Evaluation covers trajectory, not only the final sentence. See AI Agents and Tool Calling.

Production agent + RAG

The same runtime composes both data planes: retrieval as a tool or as a pre-step, models behind routing, eval on both retrieval and agent traces, observability across the whole tree.

This is the pattern people draw on whiteboards and then over-build. You need it when the product actually retrieves and acts. You do not need it because a reference architecture included both boxes.

Diagram: Four composition patterns

flowchart TB
    subgraph simple [Simple]
      S1[App] --> S2[Model]
      S1 --> S3[Eval / Obs]
    end
    subgraph rag [RAG]
      R1[App] --> R2[Gateway]
      R2 --> R3[Model]
      R1 --> R4[Retrieval]
    end
    subgraph agent [Agent]
      A1[App] --> A2[Gateway]
      A2 --> A3[Model]
      A1 --> A4[Tools / State]
    end
    subgraph combo [Agent plus RAG]
      C1[App] --> C2[Gateway]
      C2 --> C3[Model]
      C1 --> C4[Retrieval]
      C1 --> C5[Agent / Tools]
    end

Add retrieval or agents when the product requires them; do not stack patterns for completeness.

Step-by-Step Flow

A request through a composed stack (RAG + optional tools) looks like this. A simple app skips the middle hops.

Diagram: Request path through a composed stack

sequenceDiagram
    participant U as Client
    participant A as Runtime
    participant G as Gateway
    participant R as Retrieval
    participant M as Model
    participant T as Tools

    U->>A: Request
    A->>A: Auth, trace, budget
    alt needs grounding
        A->>R: Search with ACL
        R-->>A: Evidence
    end
    A->>G: Complete or stream
    G->>M: Routed model
    M-->>G: Tokens or tool call
    alt tool call
        G-->>A: Tool request
        A->>T: Execute with policy
        T-->>A: Result
        A->>G: Continue
    end
    G-->>A: Output
    A->>A: Validate, log, eval sample
    A-->>U: Response

Auth, ACLs, and tool policy stay in the runtime. The model never becomes the security boundary.

  1. Runtime authenticates, attaches trace_id, enforces payload limits.
  2. Routing policy (config or classifier) selects model tier and whether retrieval or tools are in play.
  3. Retrieval, if required, runs with tenant filters and a timeout.
  4. Gateway (or SDK wrapper) calls the model with retries and token caps.
  5. Tools, if requested, execute only if allowed; results return to the model under a step cap.
  6. Post-checks apply output policy; the runtime streams or returns.
  7. Async work appends to logs, cost counters, and sampled eval datasets.

Real Production Example

A support product starts as a single-model rewriter, then grows RAG, then one read-only tool. The code below is not a framework — it is how composition looks when layers are flags and interfaces, not a platform rewrite.

This is illustrative pseudocode: provider-specific message formats, timeout handling, tracing, retries, validation, and persistence are omitted for clarity.

from __future__ import annotations

from dataclasses import dataclass
from typing import Protocol


@dataclass(frozen=True)
class StackConfig:
    enable_retrieval: bool
    enable_tools: bool
    default_model: str
    complex_model: str
    max_tool_steps: int = 4


class Retriever(Protocol):
    async def search(self, query: str, *, tenant: str) -> list[dict]: ...


class Models(Protocol):
    async def complete(self, model: str, prompt: str, tools: list[dict] | None = None) -> dict: ...


class Tools(Protocol):
    def allowed(self, name: str, tenant: str) -> bool: ...
    def schemas(self) -> list[dict]: ...
    async def run(self, name: str, args: dict) -> str: ...


class SupportStack:
    def __init__(self, cfg: StackConfig, models: Models, retriever: Retriever | None, tools: Tools | None):
        self.cfg = cfg
        self.models = models
        self.retriever = retriever
        self.tools = tools

    def route(self, query: str, *, hard: bool) -> str:
        return self.cfg.complex_model if hard else self.cfg.default_model

    async def handle(self, query: str, tenant: str, *, hard: bool = False) -> str:
        evidence = ""
        if self.cfg.enable_retrieval:
            if self.retriever is None:
                raise RuntimeError("retrieval enabled but no retriever")
            hits = await self.retriever.search(query, tenant=tenant)
            evidence = "\n".join(h.get("text", "") for h in hits[:5])

        model = self.route(query, hard=hard)
        prompt = f"Evidence:\n{evidence}\n\nQuestion: {query}" if evidence else query
        tool_schemas = self.tools.schemas() if (self.cfg.enable_tools and self.tools) else None

        steps = 0
        messages_prompt = prompt
        while True:
            result = await self.models.complete(model, messages_prompt, tools=tool_schemas)
            call = result.get("tool_call")
            if not call or not self.cfg.enable_tools:
                return result["text"]
            if steps >= self.cfg.max_tool_steps:
                return result.get("text") or "I could not finish this request."
            if not self.tools.allowed(call["name"], tenant):
                raise PermissionError(call["name"])
            observation = await self.tools.run(call["name"], call["args"])
            messages_prompt += f"\nTool {call['name']}: {observation}"
            steps += 1

What matters: retrieval and tools are optional. Routing is a function. The model client (Models) can be a provider SDK or a gateway. Caps and ACLs are in the application. You can ship enable_retrieval=False, enable_tools=False and turn flags on when eval says the product needs them.

Design Decisions

Decision Lean choice Heavier choice Choose lean when Choose heavier when
Model access Provider SDK + retries Shared AI gateway One provider, one team Many services, keys, or failovers
Routing One pinned model Task/cost/latency policy Uniform workload Mixed easy/hard traffic
Knowledge Parametric only RAG / hybrid / graph No private corpus Cite, freshness, or private docs
Control flow Single call or fixed pipeline Agent loop Path is known Tool choice must be dynamic
Eval Spot checks Golden set in CI Throwaway prototype You will change prompts/models
Observability Structured logs Traces across hops Local debugging Production incidents and cost
Runtime Modular monolith Split retrieval/index/API One team, modest QPS Divergent scale or ownership

Decision Trade-off

Extract a gateway or a retrieval service only when copying policy across codebases is already hurting you. Premature extraction turns a composition problem into a distributed-systems problem.

Comparisons

Shape Best at Weak at Next read
Simple generation Speed, cost, clarity Private knowledge, actions LLM Concepts
RAG pipeline Grounded Q&A Open-ended side effects RAG, Enterprise RAG Architecture
Agent loop Tool use, unknown paths Predictable latency/cost AI Agents, Workflows vs Agents
Full composed stack Products that retrieve and act Small teams, early PMF AI System Architecture
This guide AI System Architecture
Capability map and composition Layered platform, contracts, ingestion vs query
When to add a box How to operate all the boxes together
Prevents overengineering Assumes you are building a serious platform

Common Mistakes

  1. Equating “production” with “every layer.” Production means operable, not maximal. A traced single-model API can be more production-ready than an unmeasured agent+RAG maze.

  2. Putting authorization in the prompt. Tenant filters belong in retrieval queries and tool allow-lists. See AI Security and Metadata Filtering.

  3. Buying a gateway to postpone naming a model ID. A gateway without a routing policy and eval is another network hop.

  4. RAG as default. If the model already knows the domain and you have no corpus, retrieval adds latency and hallucination-from-junk. Measure retrieval evaluation before adding stages.

  5. Agents as default. If the workflow is a flowchart, implement the flowchart. Workflows vs Agents.

  6. Eval as a launch-week spreadsheet. If you cannot rerun the set after a prompt change, you will regress.

  7. Logs without traces. “The LLM is slow” is not an alert. Per-hop latency is.

  8. Framework as architecture. LangChain, LangGraph, and LlamaIndex are useful libraries. They do not decide whether you need retrieval or a gateway.

Where It Breaks Down

Multi-modal and streaming products do not fit a retrieve-then-generate box. Treat speech, images, and live tools as additional model roles and I/O paths, still behind the same runtime, policy, and traces.

Organization boundaries split the stack across teams (search owns retrieval, ML owns models, backend owns the API). Without shared trace_id and version fields, composition becomes finger-pointing. The stack map is the contract; the org chart is not.

Regulatory workloads may require human review, retention rules, and region pinning that no model API will infer. Those constraints attach to the runtime and gateway policy, not to a cleverer prompt.

Research prototypes need none of this map. Forcing a stack onto a paper reproduction wastes time.

Choosing What You Actually Need

This is the section that should prevent a six-month platform build.

Decision tree: which capabilities to add

flowchart TD
    Start[What does the product do?] --> Know{Need private or current knowledge?}
    Know -->|Yes| RAG[Add retrieval]
    Know -->|No| Tools{Need tools or unknown steps?}
    RAG --> Tools
    Tools -->|Yes| Ag[Add bounded agent]
    Tools -->|No| Multi{Multiple models or providers?}
    Ag --> Multi
    Multi -->|Yes| Gw[Gateway and routing pay off]
    Multi -->|No| Risk{High risk or rapid change?}
    Gw --> Risk
    Risk -->|Yes| Ops[Stronger eval and traces]
    Risk -->|No| Lean[Model plus app plus basic obs]

Walk the questions in order. Stop when the product’s failure modes are covered.

Situation What to actually run
Simple generation Model + application + basic observability
Grounded answers over your data Add retrieval / search (hybrid and rerank as needed)
Multi-model or multi-provider production Gateway and routing become useful
Open-ended actions Add orchestration, state, tool controls
High-risk or rapidly changing system Stronger evaluation and observability
Multi-hop structured knowledge Graph / KG patterns — not more of the same vector index

The production AI stack is a menu. Order what you will operate. Leave the rest off the plate.

When NOT to Use a Full Stack

Do not stand up gateway, retrieval, agents, and a full eval platform when:

  • You are validating a single-model UX with one team and no private corpus
  • The job is offline batch generation with no interactive SLO
  • You cannot describe the failure mode each new layer would prevent
  • You do not yet have a golden set small enough to run after every prompt change — fix that before adding agents

Ship a modular monolith with clear functions. Promote a capability to a service when load, ownership, or policy duplication demands it.

Warning

A reference architecture with eight boxes is a map of the industry, not a backlog. Copying it into Kubernetes does not make the product production-grade.

Running in Production

Best Practice

Pin model, prompt, and index versions on every request. Bound every external call. Gate deploys on a golden set. Attribute tokens and errors to a route, not to “the AI.”

Dimension Guidance
Scaling Scale the API independently from indexing jobs and from retrieval replicas
Latency Budget per hop; stream generation; skip unused layers
Cost Route on task; cache where correctness allows; cap agent steps
Monitoring Trace ID across runtime → retrieval → model → tools
Evaluation CI on golden set; sample production into the set
Security Secrets in a vault; ACLs in queries and tools
Change Canary a model ID or retriever behind routing, not a big-bang deploy

Production checklist

  • Model IDs pinned and logged (not “latest”)
  • Timeouts on provider, retrieval, and tools
  • Tenant/ACL enforced outside the prompt
  • Trace ID on every hop
  • Token/cost attribution per route
  • Golden eval path for the layers you actually use
  • Rollback for prompt, index, and model independently

Interview Questions

  1. What does a production AI system need that a model API does not provide?
    Application runtime, policy (authz, retries, caps), optional retrieval and tools, evaluation, and traces — the model only emits tokens.

  2. When is an AI gateway worth extracting?
    When multiple services need shared credentials, provider failover, or spend limits. One app with one provider can wrap the SDK.

  3. How is routing different from a gateway?
    Gateway is the adapter; routing is the policy that chooses a model. Either can exist without the other.

  4. When do you add RAG?
    When answers require private, current, or citable knowledge the weights do not hold — and you can evaluate retrieval.

  5. When is an agent the wrong control flow?
    When the path is a known workflow. Use a pipeline or DAG; add a loop only when tool choice must be dynamic.

  6. How do evaluation and observability differ?
    Evaluation measures quality against a set (often offline/CI). Observability reconstructs a live request.

  7. What is the smallest production stack?
    A runtime that pins a model, times out, traces the call, and has a tiny golden set. Everything else is earned.

  8. How do you avoid overengineering?
    Add a layer only to address a named failure mode; keep capabilities as modules until they must be services.

This guide is the capability map. AI System Architecture is the platform blueprint. Specialized architectures deepen one knowledge plane.

Architecture:

Models and application behavior:

Retrieval:

Agents:

Operations:

Hands-on labs (layer illustrations, not 1:1 with this guide):

Rankings: Best AI APIs · Best Vector Databases · Best AI Agent Frameworks

Research: Foundation Models · Top Open-Source AI Models 2026 · Top AI Agent GitHub Repositories 2026

Learning path: Become an AI Engineer

Diagram: Recommended reading around this guide

flowchart LR
    LLM[LLM Concepts] --> Stack[Production AI Stack]
    Stack --> Arch[System Architecture]
    Stack --> RAG[RAG / Retrieval]
    Stack --> Ag[Agents]
    Stack --> Ops[Eval and Obs]
    Arch --> ERAG[Enterprise RAG]

Start with the stack map; go deep on the layers you actually adopt.

Learning Path

Prerequisites: Large Language Models

Next topics: AI System Architecture · AI Gateway · Model Routing · AI Copilot Architecture · Enterprise RAG Architecture · Evaluation · Observability

Estimated time: 40 min · Difficulty: Advanced

Architecture Series

Specialized production architectures for retrieval and knowledge-graph systems.

FAQs

What is a production AI stack?

The set of application capabilities around a model: access and routing, optional retrieval and agents, evaluation, observability, and the runtime that hosts them. It is a map of roles, not a required bill of materials.

Is a production AI stack the same as MLOps?

Overlaps exist (deploy, monitor, version), but this stack is about LLM applications: prompts, retrieval, tool loops, and token cost. Training pipelines and feature stores are a different discipline unless you are fine-tuning.

Do I need an AI gateway?

Only when a uniform provider interface, shared credentials, central spend limits, or failover would otherwise be duplicated. One service and one provider can call the SDK with retries.

Do I need model routing?

You always have a route — even if it is “all traffic to one ID.” Explicit routing matters when tasks differ in difficulty, latency, or cost, and you can evaluate each route. See Model Routing.

Do I need RAG?

When the answer depends on knowledge that is private, changing, or must be cited. Otherwise you are paying retrieval latency to stuff irrelevant chunks into the prompt.

Do I need agents?

When the next action is not a fixed pipeline and the system must use tools. If you can draw the flowchart, implement the flowchart.

How is this different from AI System Architecture?

This guide is composition: which capabilities you need and how they combine. AI System Architecture is how to layer a full platform (ingestion, orchestration, contracts, operations).

What should I observe first?

Model ID, latency, errors, and token counts on the request path. Add retrieval and tool spans when those layers exist. Traces beat dashboards you never query.

How do I start without overbuilding?

Pin a model, wrap the call, emit a trace, add a tiny golden set. Add retrieval when grounding fails. Add a gateway when provider policy is copied three times. Add agents when a workflow engine is the wrong shape.

Where do Labs fit?

Labs illustrate individual techniques (RAG stages, tool loops, eval). They are not a second architecture track and are not 1:1 with this page.

Which rankings are relevant?

Best AI APIs for hosted model access, Best Vector Databases if you have a retrieval plane, Best AI Agent Frameworks if you have an agent loop. Rankings do not tell you whether you need those planes.

Can I run this as a monolith?

Yes. Most teams should, until retrieval load, indexing, or ownership splits from the API. Interfaces first, services later.

References

Further Reading

Key Takeaways

  • Production AI is an application with model I/O, not a model with a UI.
  • The stack is a set of capabilities: models, gateway, routing, retrieval, agents, evaluation, observability, runtime.
  • Adopt a layer to address a failure mode; do not install the industry diagram.
  • Gateway and routing are policies around model access — often in-process at first.
  • RAG and agents are optional control and data planes with their own eval.
  • Evaluation and observability are how the rest of the stack stays honest.
  • AI System Architecture is the next read when you are ready to draw platform boundaries.

Next Topics

Learning Path

Continue Learning

Related Guides

Related Tools

ToolCategoryPurposeWebsiteBest For
LangChain
PopularOpen SourceAPI
frameworksFramework for building LLM-powered applications and workflows.langchain.comRAG systems
LangGraph
FeaturedOpen SourceAPI
frameworksGraph-based orchestration runtime for long-running, stateful agents.langgraph.devMulti-agent orchestration
LlamaIndex
Open SourceAPI
frameworksData framework for connecting LLMs to private and structured data.llamaindex.aiRAG over documents

Related Rankings