Production AI Stack in One Sentence
Production AI Stack =
- Models
- AI Gateway
- Model Routing
- RAG / Retrieval
- Agents
- Evaluation
- Observability
- Production Runtime
TL;DR
-
A production AI application is software that calls models, not a model with a UI. Prompts, weights, and a chat box do not give you retries, authorization, grounding, evaluation, or an incident trail.
-
The stack is a set of capabilities, not a shopping list. Simple generation may need only a model, an application, and basic tracing. Retrieval, gateways, agents, and heavy eval are responses to specific problems.
-
Layers are architectural roles, not required services. An AI gateway can be a library, a proxy, or a few functions in your API. Model routing can live in application code. Agents are an orchestration pattern, not a separate product you must buy.
-
Compose along failure modes. Private or changing knowledge → retrieval. Multiple models or providers → gateway and routing. Open-ended tool use → agents with bounds. High-risk or fast-changing behavior → stronger evaluation and observability.
-
Read this as a map, then go deep elsewhere. Layered platform design lives in AI System Architecture. The provider-access boundary is AI Gateway. RAG, agents, eval, and observability each have dedicated guides — this document explains how those pieces fit and when they are worth the cost.
Architecture Snapshot
Complexity
★★★★☆
Audience
AI Engineers, Tech Leads, Architects
Difficulty
Advanced
Typical Deployment
Production application
Typical Latency
Workload-dependent
Scalability
Add layers as failure modes appear
Availability Target
Match the surrounding product
Read Time
~40 min
Last Updated
September 2, 2026
Recommended Stack
- Application runtime (API + jobs)
- Model access (SDK or gateway)
- Retrieval / search when grounding is required
- Bounded agent/tool loop when paths are open-ended
- Golden-set eval + request traces
Why This Matters
A prototype is usually one call: serialize a prompt, hit a provider, return text. That shape hides most of the work that shows up the week after launch.
Production traffic is concurrent. Knowledge goes stale. Providers return 429s and 503s. The same question must sometimes hit a cheap model and sometimes a slower one. Answers that sounded fluent in a demo fail a compliance review. A customer reports a wrong answer from Tuesday and nobody can reconstruct the prompt, the retrieved chunks, or the tool calls.
Those are application problems. The model does not own retries, tenant isolation, citation checks, or cost attribution. If you do not name the capabilities around the model, they accumulate as glue in request handlers — untested, unobserved, and expensive to change.
This guide answers a practical question: what does a production AI system actually need beyond a model and a prompt? It is a composition map. The companion AI System Architecture guide is the layered platform blueprint (ingestion, orchestration, contracts, deployment). Use this page to decide which capabilities you need; use that page when you are ready to draw service boundaries.
The Problem a Production Stack Solves
Three failure modes appear when teams ship the prototype unchanged.
The model is treated as the system. Business logic, retrieval, tool permissions, and output policy sit in prompt strings. You cannot test them independently, you cannot scale them independently, and you cannot tell which part broke.
Capabilities are copied from a vendor diagram. A “full stack” appears: vector database, agent framework, gateway, eval SaaS, feature store. Most of it is idle. Complexity shows up as latency, cost, and incident surface before it shows up as quality.
There is no place to put production concerns. Retries live in one helper, cost caps in another, tracing in a third. When a request is slow, you cannot say whether retrieval, the provider, or a tool hung.
A production stack names the roles — model access, retrieval, orchestration, evaluation, observability, runtime — so you can adopt them in order, keep them thin, and avoid pretending every application is an enterprise platform.
| Without a stack map | With a stack map |
|---|---|
| Prompt file + SDK call is “the architecture” | Each capability has a job and a reason to exist |
| Every new feature adds another vendor | You add a layer when a failure mode appears |
| Incidents start at “the LLM is wrong” | You can attribute retrieval, routing, tools, or runtime |
| Eval is a spreadsheet after launch | Quality gates sit next to deploys |
How We Got Here
Production AI did not arrive as a new category of infrastructure. It is ordinary application architecture with a probabilistic component in the request path.
Diagram: A simplified evolution of production AI architectures
timeline
title Simplified architectural evolution
2020-2022 : Provider APIs
: Prompt in the handler
2023 : Chat UIs over one model
: RAG demos and vector DBs
2024 : Gateways and eval harnesses
: Tool-calling agents
2025-2026 : Compose only what you need
: Traces, golden sets, bounds
A simplified teaching sequence — not a literal industry chronology.
| Era | What shipped | What broke in production |
|---|---|---|
| Single API call | Chat wrapper over one provider | No grounding, no failover |
| RAG prototypes | Embed, retrieve, stuff context | Authz, freshness, retrieval quality |
| Framework chains | Fast composition | Opaque control flow, buried cost |
| Agent demos | Tools in a loop | Unbounded steps, irreversible actions |
| Operable systems | Routing, traces, golden sets, bounds | Requires deliberate composition |
The useful lesson is not “collect every box.” It is that models are interchangeable; the application around them is not. Gateway products, vector databases, and agent runtimes are implementations of roles. Your job is to decide which roles you have earned.
What Is the Production AI Stack?
The production AI stack is the set of capabilities an application uses to call models safely, ground them when needed, take actions when needed, and remain operable after deploy.
It is not a reference cluster you must stand up. It is not a ranking of products. It is a vocabulary for the work around the model:
- Models — what generates or classifies
- AI gateway — a common way to reach providers
- Model routing — which model handles this request
- RAG / retrieval — how external knowledge enters the prompt
- Agents — how the system chooses tools and next steps
- Evaluation — how you know quality did not regress
- Observability — how you reconstruct a request
- Production runtime — the application that hosts all of the above
Two distinctions matter.
Capability versus service. Retrieval can be a PostgreSQL extension in the same process as the API. An “AI gateway” can be a 200-line proxy or a provider SDK with retries. Agents can be a for loop with a tool table. Do not start by drawing microservices.
This stack versus AI System Architecture. This guide is what to compose. That guide is how to layer a full platform — ingestion versus query, orchestration as control plane, contracts, and operational boundaries. Specialized production patterns (Enterprise RAG, GraphRAG, Knowledge Graph + LLM) sit on top of both.
The Production AI Stack
At a glance, a production AI application is a runtime that may use some or all of the following. Nothing below the application is mandatory.
| Layer | Question it answers | You need it when |
|---|---|---|
| Models | What can generate, classify, embed, or extract? | Always — this is the non-negotiable core |
| AI gateway | How do we talk to providers uniformly? | Multiple providers, shared credentials, or central policy |
| Model routing | Which model (or tier) for this request? | Quality, cost, latency, or availability differ by task |
| RAG / retrieval | What evidence is not in the weights? | Private, current, or citable knowledge |
| Agents | What should we do next, including side effects? | The control flow is not a fixed pipeline |
| Evaluation | Did this change make answers worse? | You ship prompts, retrieval, or tools more than once |
| Observability | What happened on request X? | You have users, SLOs, or a cost bill |
| Production runtime | Where does this run, persist, and fail? | Always — this is still an application |
Gateway and routing overlap: routing is often a policy the gateway executes. They are listed separately because a single-provider app can route between model tiers without a gateway, and a gateway can exist solely for credentials and retries with one model behind it.
Architecture
Diagram: Production AI stack
flowchart TB
Runtime[Production Runtime]
Runtime --> Models
Runtime --> Gateway[AI Gateway]
Gateway --> Routing[Model Routing]
Routing --> Models
Runtime --> RAG[RAG / Retrieval]
RAG --> Search[Vector / Search]
Runtime --> Agents
Agents --> Tools[Tools / Memory]
Runtime --> Eval[Evaluation]
Runtime --> Obs[Observability]
The runtime is the application. Other boxes are capabilities it may call — not a required topology.
The diagram is a map of roles. Edges are “may use,” not “must deploy.” A simple generator is Runtime → Models, with observability as logs and traces on that path. A RAG assistant adds retrieval. An agent assistant adds a bounded loop over tools. Evaluation is usually offline-plus-CI, not a box on the hot path — except when you sample production traces into an eval set.
Engineering Insight
If you cannot explain a layer as a failure mode you have already hit (or will hit this quarter), it is decoration. Architecture diagrams are not procurement checklists.
1. Models
The model layer is the foundation: one or more models that generate text, classify intent, embed queries, extract structure, or score candidates.
Production applications rarely use a single model for everything. Embedding models are not chat models. A cheap classifier that routes “reset password” away from a frontier model is often a better architecture than a smarter prompt. Vision, speech, and rerankers are additional model roles, not features of “the LLM.”
Selection is an engineering decision, not a leaderboard decision. You care about:
- Capability — reasoning, tool calling, long context, structured output, multilingual
- Cost — tokens in and out, embedding volume, rerank calls
- Latency — time to first token and time to last token under your concurrency
- Context — window size versus the documents and history you actually send
- Control — data handling, region, rate limits, deprecation policy
Large Language Models covers how these systems work. Family guides (GPT, Claude, Gemini, Llama, Mistral, DeepSeek) are for model-specific behavior. Tokens and Context Windows are the budgets the rest of the stack must respect. Cost Optimization and Latency Optimization are how this layer is operated.
Hosted access is compared independently in Best AI APIs. That ranking is about APIs, not about whether you need a gateway.
Decision Trade-off
One strong model simplifies operations and eval. Multiple models reduce unit cost and tail latency if — and only if — you can route correctly and measure quality per route. Routing without eval is a cost optimization that silently degrades answers.
You do not need a model garden on day one. You need a pinned model ID, a way to change it without rewriting handlers, and an eval set that tells you the change was not a regression.
2. AI Gateway
An AI gateway is a common interface in front of model providers. Architecturally it is an adapter plus policy: one request shape, many backends.
Typical responsibilities, when you actually need them:
- Uniform API — chat completions, embeddings, and (sometimes) rerank behind one client
- Credentials, quotas, and policy — centralize these when multiple services need the same provider controls
- Retries and fallbacks — 429/503 handling, secondary provider, and, where duplicate work/cost is acceptable, hedged requests
- Usage and cost controls — per-key budgets, max tokens, block lists
- Telemetry — one place to emit provider latency, status, and token counts
A gateway is not mandatory. A single-provider application with one model and moderate traffic can call the provider SDK, wrap retries, and emit traces from the application. The gateway earns its keep when provider diversity, shared credentials, or central policy would otherwise be copied into every service.
Do not confuse an AI gateway with your public API gateway. The public gateway authenticates users and rate-limits clients. An AI gateway is the adapter toward model providers — it can centralize provider credentials and spend controls when those belong in one place. They can be the same process; they are different jobs.
This is not a product guide. Implementations range from a thin internal proxy to libraries such as LiteLLM. If you only need retries and a second model ID, write that in the application. If ten services all need the same provider failover, extract a gateway.
The dedicated AI Gateway guide covers when that boundary is worth introducing, what belongs there, and what does not. Related operational concerns also live in Cost Optimization, AI Security, and Observability.
3. Model Routing
Model routing is the policy that sends different requests to different models. It is often implemented inside a gateway, but it is a separate idea: the gateway is how you call models; routing is which model you call.
Applications route on:
- Task — classify, extract, generate, embed, rerank
- Quality — hard reasoning versus template filling
- Latency — interactive chat versus batch
- Cost — default cheap, escalate on low confidence
- Availability — failover when a provider is dark
- Context — a long document needs a long-context tier; a one-line FAQ does not
Routing can be a static map (intent → model_id), a small classifier, or a cascade (try cheap, escalate if the eval proxy is low). The policy belongs in application or gateway config, versioned and logged, so you can answer “which model served this tenant last Thursday?”
A single-model app still “routes”: every request goes to one pinned ID. Name that pin. The day you add a second model, you already have a routing table instead of a hunt through handlers.
Warning
Routing on cost without a quality gate will eventually send the hard cases to the cheap model. Pair every new route with a slice of the evaluation set.
The dedicated Model Routing guide covers how that decision is made — from a static task map to production policies involving capability, quality, latency, cost, availability, fallback, and evaluation. Treat routing as a versioned policy next to Cost Optimization and Latency Optimization, enforced where you already emit traces.
4. RAG and Retrieval
Retrieval exists because weights are the wrong store for private, current, or citable knowledge. RAG (retrieval-augmented generation) is the pattern: fetch evidence, then generate with that evidence in context.
Architecturally, RAG is a data plane the application calls before (or during) generation:
- Turn the user question into one or more search queries (Query Transformation when naive queries fail).
- Retrieve candidates from a vector database / vector search index, often fused with lexical search (Hybrid Search).
- Enforce authorization before retrieval and carry the resulting tenant/ACL constraints into the retrieval query (Metadata Filtering) — never rely on the prompt to enforce access.
- Optionally re-rank a wider candidate set down to what fits the context window.
- Build a prompt that cites sources; generate; check faithfulness.
Embeddings and Embedding Models are how text becomes searchable. Chunking Strategies decide what “a document” means to the index. None of this replaces the model; it changes what the model is allowed to see.
You do not need RAG for a generator that only needs parametric knowledge (tone, format, coding assistance on public APIs). You do need it when wrong answers come from missing files, stale facts, or an inability to cite.
Quality of this layer is measurable. Retrieval Evaluation scores the index; RAG Evaluation attributes failures to retrieval versus generation. The Basic RAG, Hybrid RAG, Reranking, and Retrieval Evaluation labs show those pipelines as traces — they are not a required companion to this architecture guide.
When RAG must survive tenants, ACLs, and SLOs, the production envelope is Enterprise RAG Architecture. When answers need multi-hop structure, GraphRAG / GraphRAG Architecture and Knowledge Graph + LLM are different knowledge planes, not “more vectors.”
Compare stores independently in Best Vector Databases only after you know you need a retrieval plane.
5. Agents
An agent is an orchestration pattern: the model proposes the next action, the application executes it, and the loop continues until a stop condition. A RAG pipeline is usually a fixed retrieve-then-generate path. You add agentic control when the sequence of steps is not known in advance and the system must use tools, branch, or retry.
AI Agents is the concept guide. Architecturally, agents add:
- Tool calling / function calling — the model names a tool and arguments; your code validates, authorizes, and executes
- Workflow versus loop — Workflows vs Agents: if the path is a DAG, use a workflow; if the model must choose the path, use a bounded loop
- State — what is in the messages versus what is in a store
- Memory — facts that must survive a session (and go stale)
- Human-in-the-loop — approval gates for irreversible actions
- Durable execution — waits, retries, and resumes measured in hours, not milliseconds
- Protocol-shaped tools — Model Context Protocol standardizes tool I/O; it does not replace permissions or tracing
- Reliability — Agent Evaluation scores outcomes and trajectories separately; Guardrails constrain tools and outputs
Agents are expensive: extra model calls, extra latency variance, extra ways to duplicate side effects. Cap steps, cap tokens, require idempotency on writes, and do not let the model be the authorization layer.
Agentic RAG is the hybrid: the agent may retrieve more than once, or retrieve as a tool, instead of a single retrieve-then-generate pass. Use it when one retrieval is not enough — and measure it, because extra hops are extra failure modes.
Hands-on traces: Tool Calling, Agent Evaluation, Memory. Those labs map to their own guides; this stack page does not own a Lab.
Frameworks compared in Best AI Agent Frameworks implement the loop. They are libraries behind an application interface, not a substitute for timeouts, tool allow-lists, and eval.
Best Practice
Prefer a fixed pipeline or an explicit workflow until tool choice must be dynamic. Autonomy is a cost and a reliability tax — spend it on tasks that are actually open-ended.
6. Evaluation
Traditional software tests assert deterministic outputs. Model outputs are distributions. Evaluation is the regression suite for that reality: a versioned golden set, layered metrics, and a gate that can block a deploy.
In a production stack, evaluation is not one score. It is several questions:
| Question | Where it lives |
|---|---|
| Did the model follow the spec? | LLM Evaluation, Prompt Evaluation |
| Did we retrieve the right evidence? | Retrieval Evaluation |
| Was the answer faithful and useful? | RAG Evaluation |
| Did the agent take a sensible path? | Agent Evaluation |
| Did a public leaderboard move? | Benchmarks — shortlist only |
Regression testing means the golden set is checked in CI when you change prompts, models, chunking, or tools. Production feedback means sampled traces, thumbs-down, and corrections flow into that set — otherwise you are testing yesterday’s product.
Evaluation is how you know a routing change, a new retriever, or a cheaper model did not quietly fail. Without it, every other layer is an untested hypothesis. Hallucination Detection and Guardrails are complementary: they catch some failures online; eval tells you the system got worse before users do.
You do not need a research-grade harness for a weekend prototype. You need one the moment you will change the system more than once and still claim quality.
7. Observability
Observability is request-scoped visibility: what you need to reconstruct and attribute a failure. It is not the same as evaluation. Eval answers “is quality acceptable?” Observability answers “what happened on this request?”
For an AI application, a useful trace includes:
- The user request — identity, tenant, route, payload size
- Model calls — model ID, prompt version, tokens in/out, finish reason
- Latency — per hop (gateway, retrieval, generation, tools), not one blob
- Errors — provider status, timeouts, validation failures
- Usage — tokens and estimated cost, attributed to tenant/route
- Retrieval — query, filters, hit count, chunk IDs (not necessarily raw text in the log store)
- Tool calls — name, arguments hash, duration, success/failure
- Parent/child spans — so an agent loop is a tree, not a pile of log lines
If you cannot answer “why was this slow?” or “why did this say X?” from a trace_id, you do not have production observability yet — you have stdout.
Keep instrumentation boring: OpenTelemetry-style traces plus structured logs. Product-specific UIs are optional. Do not build an “AI observability platform” as a prerequisite for shipping; emit spans from day one and add backends when you have volume.
Cost Optimization and Latency Optimization depend on these signals. AI Security depends on not putting secrets and raw PII into the trace store.
8. Production Runtime
The runtime is the application and infrastructure that hosts the stack. It is the part teams already know how to build — and then forget when the demo is a notebook.
At architecture level:
- APIs and services — the synchronous request path (HTTP, RPC, streaming)
- Queues and jobs — indexing, eval batches, long agent runs, webhook fan-out
- Persistence — sessions, audit logs, documents, index versions, golden sets
- Caching — exact and semantic caching; see Caching
- Secrets — provider keys, tool credentials; never in prompts or client bundles
- Scaling — stateless app replicas versus retrieval replicas versus job workers
- Reliability — timeouts, bulkheads, backpressure, idempotency, deploy rollback
This layer is why “just call the model” fails: the model is one dependency among databases, queues, and identity. AI System Architecture goes deeper on ingestion versus query paths and orchestration as a control plane. AI Security belongs here as threat modeling, not as a system prompt.
If you already run a production web service, you already have most of this. The AI-specific work is pinning model/index/prompt versions on every request, bounding non-deterministic loops, and treating providers as unreliable backends.
How the Pieces Fit Together
The stack is compositional. Four patterns cover most products. None of them require every layer.
Simple AI application
Application → Model → Evaluation / Observability
A formatter, classifier, or writing aid with no private corpus and one provider. The runtime is an API. Observability is traces with model ID and token counts. Evaluation is a small golden set on format and a handful of quality cases. No gateway, no RAG, no agent loop.
RAG application
Application → Gateway → Model
↘ Retrieval → Vector / Search → Reranking
Grounded Q&A over your documents. Retrieval is on the hot path; the gateway is optional until you have more than one model or provider. Evaluation must include retrieval metrics, not only “the answer sounds good.” See RAG and Enterprise RAG Architecture.
Agent application
Application → Gateway → Model
↘ Agent → Tools / State / Memory
Open-ended work with side effects: tickets, queries, code, browser. The loop is bounded. Tools are permissioned in the application. Memory is an explicit store, not “whatever fit in the window.” Evaluation covers trajectory, not only the final sentence. See AI Agents and Tool Calling.
Production agent + RAG
The same runtime composes both data planes: retrieval as a tool or as a pre-step, models behind routing, eval on both retrieval and agent traces, observability across the whole tree.
This is the pattern people draw on whiteboards and then over-build. You need it when the product actually retrieves and acts. You do not need it because a reference architecture included both boxes.
Diagram: Four composition patterns
flowchart TB
subgraph simple [Simple]
S1[App] --> S2[Model]
S1 --> S3[Eval / Obs]
end
subgraph rag [RAG]
R1[App] --> R2[Gateway]
R2 --> R3[Model]
R1 --> R4[Retrieval]
end
subgraph agent [Agent]
A1[App] --> A2[Gateway]
A2 --> A3[Model]
A1 --> A4[Tools / State]
end
subgraph combo [Agent plus RAG]
C1[App] --> C2[Gateway]
C2 --> C3[Model]
C1 --> C4[Retrieval]
C1 --> C5[Agent / Tools]
end
Add retrieval or agents when the product requires them; do not stack patterns for completeness.
Step-by-Step Flow
A request through a composed stack (RAG + optional tools) looks like this. A simple app skips the middle hops.
Diagram: Request path through a composed stack
sequenceDiagram
participant U as Client
participant A as Runtime
participant G as Gateway
participant R as Retrieval
participant M as Model
participant T as Tools
U->>A: Request
A->>A: Auth, trace, budget
alt needs grounding
A->>R: Search with ACL
R-->>A: Evidence
end
A->>G: Complete or stream
G->>M: Routed model
M-->>G: Tokens or tool call
alt tool call
G-->>A: Tool request
A->>T: Execute with policy
T-->>A: Result
A->>G: Continue
end
G-->>A: Output
A->>A: Validate, log, eval sample
A-->>U: Response
Auth, ACLs, and tool policy stay in the runtime. The model never becomes the security boundary.
- Runtime authenticates, attaches
trace_id, enforces payload limits. - Routing policy (config or classifier) selects model tier and whether retrieval or tools are in play.
- Retrieval, if required, runs with tenant filters and a timeout.
- Gateway (or SDK wrapper) calls the model with retries and token caps.
- Tools, if requested, execute only if allowed; results return to the model under a step cap.
- Post-checks apply output policy; the runtime streams or returns.
- Async work appends to logs, cost counters, and sampled eval datasets.
Real Production Example
A support product starts as a single-model rewriter, then grows RAG, then one read-only tool. The code below is not a framework — it is how composition looks when layers are flags and interfaces, not a platform rewrite.
This is illustrative pseudocode: provider-specific message formats, timeout handling, tracing, retries, validation, and persistence are omitted for clarity.
from __future__ import annotations
from dataclasses import dataclass
from typing import Protocol
@dataclass(frozen=True)
class StackConfig:
enable_retrieval: bool
enable_tools: bool
default_model: str
complex_model: str
max_tool_steps: int = 4
class Retriever(Protocol):
async def search(self, query: str, *, tenant: str) -> list[dict]: ...
class Models(Protocol):
async def complete(self, model: str, prompt: str, tools: list[dict] | None = None) -> dict: ...
class Tools(Protocol):
def allowed(self, name: str, tenant: str) -> bool: ...
def schemas(self) -> list[dict]: ...
async def run(self, name: str, args: dict) -> str: ...
class SupportStack:
def __init__(self, cfg: StackConfig, models: Models, retriever: Retriever | None, tools: Tools | None):
self.cfg = cfg
self.models = models
self.retriever = retriever
self.tools = tools
def route(self, query: str, *, hard: bool) -> str:
return self.cfg.complex_model if hard else self.cfg.default_model
async def handle(self, query: str, tenant: str, *, hard: bool = False) -> str:
evidence = ""
if self.cfg.enable_retrieval:
if self.retriever is None:
raise RuntimeError("retrieval enabled but no retriever")
hits = await self.retriever.search(query, tenant=tenant)
evidence = "\n".join(h.get("text", "") for h in hits[:5])
model = self.route(query, hard=hard)
prompt = f"Evidence:\n{evidence}\n\nQuestion: {query}" if evidence else query
tool_schemas = self.tools.schemas() if (self.cfg.enable_tools and self.tools) else None
steps = 0
messages_prompt = prompt
while True:
result = await self.models.complete(model, messages_prompt, tools=tool_schemas)
call = result.get("tool_call")
if not call or not self.cfg.enable_tools:
return result["text"]
if steps >= self.cfg.max_tool_steps:
return result.get("text") or "I could not finish this request."
if not self.tools.allowed(call["name"], tenant):
raise PermissionError(call["name"])
observation = await self.tools.run(call["name"], call["args"])
messages_prompt += f"\nTool {call['name']}: {observation}"
steps += 1
What matters: retrieval and tools are optional. Routing is a function. The model client (Models) can be a provider SDK or a gateway. Caps and ACLs are in the application. You can ship enable_retrieval=False, enable_tools=False and turn flags on when eval says the product needs them.
Design Decisions
| Decision | Lean choice | Heavier choice | Choose lean when | Choose heavier when |
|---|---|---|---|---|
| Model access | Provider SDK + retries | Shared AI gateway | One provider, one team | Many services, keys, or failovers |
| Routing | One pinned model | Task/cost/latency policy | Uniform workload | Mixed easy/hard traffic |
| Knowledge | Parametric only | RAG / hybrid / graph | No private corpus | Cite, freshness, or private docs |
| Control flow | Single call or fixed pipeline | Agent loop | Path is known | Tool choice must be dynamic |
| Eval | Spot checks | Golden set in CI | Throwaway prototype | You will change prompts/models |
| Observability | Structured logs | Traces across hops | Local debugging | Production incidents and cost |
| Runtime | Modular monolith | Split retrieval/index/API | One team, modest QPS | Divergent scale or ownership |
Decision Trade-off
Extract a gateway or a retrieval service only when copying policy across codebases is already hurting you. Premature extraction turns a composition problem into a distributed-systems problem.
Comparisons
| Shape | Best at | Weak at | Next read |
|---|---|---|---|
| Simple generation | Speed, cost, clarity | Private knowledge, actions | LLM Concepts |
| RAG pipeline | Grounded Q&A | Open-ended side effects | RAG, Enterprise RAG Architecture |
| Agent loop | Tool use, unknown paths | Predictable latency/cost | AI Agents, Workflows vs Agents |
| Full composed stack | Products that retrieve and act | Small teams, early PMF | AI System Architecture |
| This guide | AI System Architecture |
|---|---|
| Capability map and composition | Layered platform, contracts, ingestion vs query |
| When to add a box | How to operate all the boxes together |
| Prevents overengineering | Assumes you are building a serious platform |
Common Mistakes
-
Equating “production” with “every layer.” Production means operable, not maximal. A traced single-model API can be more production-ready than an unmeasured agent+RAG maze.
-
Putting authorization in the prompt. Tenant filters belong in retrieval queries and tool allow-lists. See AI Security and Metadata Filtering.
-
Buying a gateway to postpone naming a model ID. A gateway without a routing policy and eval is another network hop.
-
RAG as default. If the model already knows the domain and you have no corpus, retrieval adds latency and hallucination-from-junk. Measure retrieval evaluation before adding stages.
-
Agents as default. If the workflow is a flowchart, implement the flowchart. Workflows vs Agents.
-
Eval as a launch-week spreadsheet. If you cannot rerun the set after a prompt change, you will regress.
-
Logs without traces. “The LLM is slow” is not an alert. Per-hop latency is.
-
Framework as architecture. LangChain, LangGraph, and LlamaIndex are useful libraries. They do not decide whether you need retrieval or a gateway.
Where It Breaks Down
Multi-modal and streaming products do not fit a retrieve-then-generate box. Treat speech, images, and live tools as additional model roles and I/O paths, still behind the same runtime, policy, and traces.
Organization boundaries split the stack across teams (search owns retrieval, ML owns models, backend owns the API). Without shared trace_id and version fields, composition becomes finger-pointing. The stack map is the contract; the org chart is not.
Regulatory workloads may require human review, retention rules, and region pinning that no model API will infer. Those constraints attach to the runtime and gateway policy, not to a cleverer prompt.
Research prototypes need none of this map. Forcing a stack onto a paper reproduction wastes time.
Choosing What You Actually Need
This is the section that should prevent a six-month platform build.
Decision tree: which capabilities to add
flowchart TD
Start[What does the product do?] --> Know{Need private or current knowledge?}
Know -->|Yes| RAG[Add retrieval]
Know -->|No| Tools{Need tools or unknown steps?}
RAG --> Tools
Tools -->|Yes| Ag[Add bounded agent]
Tools -->|No| Multi{Multiple models or providers?}
Ag --> Multi
Multi -->|Yes| Gw[Gateway and routing pay off]
Multi -->|No| Risk{High risk or rapid change?}
Gw --> Risk
Risk -->|Yes| Ops[Stronger eval and traces]
Risk -->|No| Lean[Model plus app plus basic obs]
Walk the questions in order. Stop when the product’s failure modes are covered.
| Situation | What to actually run |
|---|---|
| Simple generation | Model + application + basic observability |
| Grounded answers over your data | Add retrieval / search (hybrid and rerank as needed) |
| Multi-model or multi-provider production | Gateway and routing become useful |
| Open-ended actions | Add orchestration, state, tool controls |
| High-risk or rapidly changing system | Stronger evaluation and observability |
| Multi-hop structured knowledge | Graph / KG patterns — not more of the same vector index |
The production AI stack is a menu. Order what you will operate. Leave the rest off the plate.
When NOT to Use a Full Stack
Do not stand up gateway, retrieval, agents, and a full eval platform when:
- You are validating a single-model UX with one team and no private corpus
- The job is offline batch generation with no interactive SLO
- You cannot describe the failure mode each new layer would prevent
- You do not yet have a golden set small enough to run after every prompt change — fix that before adding agents
Ship a modular monolith with clear functions. Promote a capability to a service when load, ownership, or policy duplication demands it.
Warning
A reference architecture with eight boxes is a map of the industry, not a backlog. Copying it into Kubernetes does not make the product production-grade.
Running in Production
Best Practice
Pin model, prompt, and index versions on every request. Bound every external call. Gate deploys on a golden set. Attribute tokens and errors to a route, not to “the AI.”
| Dimension | Guidance |
|---|---|
| Scaling | Scale the API independently from indexing jobs and from retrieval replicas |
| Latency | Budget per hop; stream generation; skip unused layers |
| Cost | Route on task; cache where correctness allows; cap agent steps |
| Monitoring | Trace ID across runtime → retrieval → model → tools |
| Evaluation | CI on golden set; sample production into the set |
| Security | Secrets in a vault; ACLs in queries and tools |
| Change | Canary a model ID or retriever behind routing, not a big-bang deploy |
Production checklist
- Model IDs pinned and logged (not “latest”)
- Timeouts on provider, retrieval, and tools
- Tenant/ACL enforced outside the prompt
- Trace ID on every hop
- Token/cost attribution per route
- Golden eval path for the layers you actually use
- Rollback for prompt, index, and model independently
Interview Questions
-
What does a production AI system need that a model API does not provide?
Application runtime, policy (authz, retries, caps), optional retrieval and tools, evaluation, and traces — the model only emits tokens. -
When is an AI gateway worth extracting?
When multiple services need shared credentials, provider failover, or spend limits. One app with one provider can wrap the SDK. -
How is routing different from a gateway?
Gateway is the adapter; routing is the policy that chooses a model. Either can exist without the other. -
When do you add RAG?
When answers require private, current, or citable knowledge the weights do not hold — and you can evaluate retrieval. -
When is an agent the wrong control flow?
When the path is a known workflow. Use a pipeline or DAG; add a loop only when tool choice must be dynamic. -
How do evaluation and observability differ?
Evaluation measures quality against a set (often offline/CI). Observability reconstructs a live request. -
What is the smallest production stack?
A runtime that pins a model, times out, traces the call, and has a tiny golden set. Everything else is earned. -
How do you avoid overengineering?
Add a layer only to address a named failure mode; keep capabilities as modules until they must be services.
Related Guides
This guide is the capability map. AI System Architecture is the platform blueprint. Specialized architectures deepen one knowledge plane.
Architecture:
- AI System Architecture
- AI Gateway
- Model Routing
- AI Copilot Architecture
- Enterprise RAG Architecture
- GraphRAG Architecture
- Knowledge Graph + LLM Architecture
- Enterprise Knowledge Graph Architecture
Models and application behavior:
Retrieval:
Agents:
- AI Agents · Tool Calling · Workflows vs Agents · Memory · Human-in-the-Loop · Agent Evaluation · Model Context Protocol
Operations:
- Evaluation · Observability · Cost Optimization · Latency Optimization · Caching · AI Security · Guardrails
Hands-on labs (layer illustrations, not 1:1 with this guide):
Rankings: Best AI APIs · Best Vector Databases · Best AI Agent Frameworks
Research: Foundation Models · Top Open-Source AI Models 2026 · Top AI Agent GitHub Repositories 2026
Learning path: Become an AI Engineer
Diagram: Recommended reading around this guide
flowchart LR
LLM[LLM Concepts] --> Stack[Production AI Stack]
Stack --> Arch[System Architecture]
Stack --> RAG[RAG / Retrieval]
Stack --> Ag[Agents]
Stack --> Ops[Eval and Obs]
Arch --> ERAG[Enterprise RAG]
Start with the stack map; go deep on the layers you actually adopt.
Learning Path
Prerequisites: Large Language Models
Next topics: AI System Architecture · AI Gateway · Model Routing · AI Copilot Architecture · Enterprise RAG Architecture · Evaluation · Observability
Estimated time: 40 min · Difficulty: Advanced
Architecture Series
Specialized production architectures for retrieval and knowledge-graph systems.
FAQs
What is a production AI stack?
The set of application capabilities around a model: access and routing, optional retrieval and agents, evaluation, observability, and the runtime that hosts them. It is a map of roles, not a required bill of materials.
Is a production AI stack the same as MLOps?
Overlaps exist (deploy, monitor, version), but this stack is about LLM applications: prompts, retrieval, tool loops, and token cost. Training pipelines and feature stores are a different discipline unless you are fine-tuning.
Do I need an AI gateway?
Only when a uniform provider interface, shared credentials, central spend limits, or failover would otherwise be duplicated. One service and one provider can call the SDK with retries.
Do I need model routing?
You always have a route — even if it is “all traffic to one ID.” Explicit routing matters when tasks differ in difficulty, latency, or cost, and you can evaluate each route. See Model Routing.
Do I need RAG?
When the answer depends on knowledge that is private, changing, or must be cited. Otherwise you are paying retrieval latency to stuff irrelevant chunks into the prompt.
Do I need agents?
When the next action is not a fixed pipeline and the system must use tools. If you can draw the flowchart, implement the flowchart.
How is this different from AI System Architecture?
This guide is composition: which capabilities you need and how they combine. AI System Architecture is how to layer a full platform (ingestion, orchestration, contracts, operations).
What should I observe first?
Model ID, latency, errors, and token counts on the request path. Add retrieval and tool spans when those layers exist. Traces beat dashboards you never query.
How do I start without overbuilding?
Pin a model, wrap the call, emit a trace, add a tiny golden set. Add retrieval when grounding fails. Add a gateway when provider policy is copied three times. Add agents when a workflow engine is the wrong shape.
Where do Labs fit?
Labs illustrate individual techniques (RAG stages, tool loops, eval). They are not a second architecture track and are not 1:1 with this page.
Which rankings are relevant?
Best AI APIs for hosted model access, Best Vector Databases if you have a retrieval plane, Best AI Agent Frameworks if you have an agent loop. Rankings do not tell you whether you need those planes.
Can I run this as a monolith?
Yes. Most teams should, until retrieval load, indexing, or ownership splits from the API. Interfaces first, services later.
References
- Anthropic — Building effective agents
- OpenTelemetry Documentation
- OpenAI API Documentation
- Model Context Protocol Specification
- Microsoft Azure GPT-RAG Solution Accelerator
Further Reading
- Anthropic Documentation
- Google AI for Developers
- LiteLLM Documentation
- LangChain Documentation
- LlamaIndex Documentation
Key Takeaways
- Production AI is an application with model I/O, not a model with a UI.
- The stack is a set of capabilities: models, gateway, routing, retrieval, agents, evaluation, observability, runtime.
- Adopt a layer to address a failure mode; do not install the industry diagram.
- Gateway and routing are policies around model access — often in-process at first.
- RAG and agents are optional control and data planes with their own eval.
- Evaluation and observability are how the rest of the stack stays honest.
- AI System Architecture is the next read when you are ready to draw platform boundaries.