AI Gateway in One Sentence
AI Gateway =
- Provider adapters
- Credentials and policy
- Usage telemetry
- Optional routing
TL;DR
-
An AI gateway is a runtime boundary, not a product category you must buy. It sits between application services and model providers so credentials, policy, retries, and usage accounting are not copied into every caller.
-
It exists because provider integrations multiply. Each new service that talks to OpenAI, Anthropic, Google, and self-hosted models independently reimplements keys, quotas, error handling, and telemetry — until that duplication becomes the operational problem.
-
A gateway is an adapter plus policy. Common jobs include provider abstraction, centralized credentials, rate limits, spend controls, request normalization, and traces. None of those jobs is mandatory for every application.
-
It is not an AI platform. Business logic, agent orchestration, RAG retrieval, evaluation, and long-term memory do not belong here. Centralizing those concerns turns the gateway into a monolith that every team has to wait on.
-
Skip it when the system is still simple. One service, one provider, and straightforward credentials usually want a provider SDK with timeouts and traces — not another hop. Production AI Stack is the composition map; this guide is the gateway boundary in detail.
Architecture Snapshot
Complexity
★★★☆☆
Audience
AI Engineers, Platform Engineers, Architects
Difficulty
Advanced
Typical Deployment
In-process client or shared proxy
Typical Latency
Adds a hop; keep it thin
Scalability
Earned by multi-service / multi-provider load
Availability Target
Meet the application's SLO; avoid making the gateway a weaker availability boundary
Read Time
~35 min
Last Updated
September 5, 2026
Recommended Starting Point
- Provider SDK + timeouts + traces (no gateway)
- Extract a gateway when credentials or policy are copied across services
- Keep routing, RAG, and agents outside the gateway unless they are provider-access policy
Why This Matters
A prototype usually calls one model from one process. The SDK lives in the request handler. The API key lives in an environment variable. That shape is correct for a long time.
The shape breaks when the second service needs the same provider, or the first service needs a second provider. Keys leak into three repositories. Each team writes a slightly different retry loop. Finance cannot tell which product spent the tokens. A 429 in one handler is a crash; in another it is a silent empty reply. Switching a model ID means hunting through services that all speak a slightly different request dialect.
Those are not model problems. They are integration-boundary problems. An AI gateway is the name for a place that owns that boundary: how the organization talks to model providers, under what policy, and with what evidence after the fact.
This guide answers a practical question: why would an AI application need an AI gateway between its services and model providers? It is not a ranking of gateway products. It is not an argument that every production system needs one. Production AI Stack tells you which capabilities to compose; AI System Architecture tells you how to layer a platform. This page is only the provider-access boundary.
The Problem an AI Gateway Solves
The operational problem is duplicated provider integration.
Imagine four services — a chatbot API, a summarization job, a classification worker, and an internal eval runner — each calling hosted APIs and maybe a self-hosted model. Independently they each acquire:
- Provider SDKs and request/response shapes that differ by vendor
- Credentials and secret rotation
- Per-key or per-org quotas and spend
- Retry, timeout, and fallback behavior
- Telemetry for latency, errors, and token counts
- Informal policy (“this tenant cannot use the expensive model”)
None of those concerns is exotic. Each is cheap to write once. They become expensive when they drift: one service retries forever, another never retries; one logs token usage, another only logs HTTP status; one still holds a deprecated key.
A gateway does not invent these jobs. It moves them to one runtime boundary so application services can call “complete this request under this tenant and policy” instead of speaking every provider dialect.
You do not need every control on day one. A two-person app with one provider has a credentials problem, not a gateway problem. The gateway earns its keep when copying the integration is already hurting you — or is about to, because several services must share keys, spend limits, or failover.
| Without a shared boundary | With an AI gateway boundary |
|---|---|
| Every service embeds provider SDKs and keys | Services call one access layer |
| Quotas and spend are guessed from invoices | Usage can be attributed per service or tenant |
| Retries and timeouts differ by handler | Retry policy is named and versioned |
| A provider outage is a per-service incident | Failover and timeouts are a shared path |
| Traces stop at “we called an LLM” | Provider, model ID, tokens, and status are on-span |
Note
Centralization is a trade. You add a hop, an operational surface, and a potential single point of failure. That is justified only when the duplication cost is already real.
How We Got Here
Production teams did not wake up wanting another proxy. The boundary appeared as provider APIs became a default dependency.
| Era | What teams did | What started to hurt |
|---|---|---|
| Single SDK call | One handler, one provider, one key | Almost nothing — this is still the right start |
| Many product surfaces | Chat, batch, eval, copilots all call models | Keys, quotas, and retry logic copied everywhere |
| Multi-provider reality | Hosted APIs plus open-source or self-hosted models | Divergent request shapes and failure modes |
| Shared platform | A common access layer with policy and traces | Overgrowth if the layer absorbs product logic |
Gateway libraries and proxies showed up as a convenient packaging of adapters: one client shape, many backends. That packaging is useful. It is not architecture by itself. A 200-line internal client with timeouts, a secrets manager, and OpenTelemetry spans is an AI gateway. A purchased proxy with no policy and no traces is just another network hop.
The useful lesson is the same one as the rest of the production AI stack: name the role, then decide how thin the implementation can be.
What Is an AI Gateway?
An AI gateway is a centralized runtime boundary that applications use to access and govern model providers through a shared adapter and policy layer.
Three words in that definition do work:
- Runtime — it sits on the request path, not in a design doc. It can be in-process (a client library) or out-of-process (a proxy/service).
- Boundary — it is the application-facing provider-access boundary (and the first hop back into the application). Application identity, tenant, and policy should be explicit before that boundary. Network infrastructure such as egress proxies or service meshes may still sit beyond it.
- Adapter and policy — it translates a stable internal request into provider-specific calls, and it can enforce organization rules on those calls.
It is not automatically:
- An API gateway for your public HTTP APIs
- A model router (routing is a decision; the gateway is a place that decision may run — see AI Gateway vs Model Routing)
- A model-serving engine (vLLM, TensorRT-LLM, and similar run models; a gateway calls them)
- An agent framework or orchestrator
- A guarantee of lower cost or stronger security
Hosted model access is compared independently in Best AI APIs. That ranking is about APIs, not about whether you need this boundary. Implementations range from a thin internal proxy to libraries such as LiteLLM. This guide does not rank those products.
How an AI Gateway Works
The simplest mental model is three boxes:
Diagram: AI gateway as a provider-access boundary
flowchart TB
App[Application / AI service]
GW[AI Gateway]
P1[Provider A]
P2[Provider B]
P3[Self-hosted model]
App --> GW
GW --> P1
GW --> P2
GW --> P3
Applications call the gateway; in this architecture, the gateway owns the provider-access boundary.
Why that extra box exists: provider APIs are not a stable internal contract. Auth headers, error codes, streaming, tool-call payloads, token accounting, and rate-limit semantics differ. If five services each absorb those differences, you have five integrations to change when a provider deprecates a model or adds a required header.
A typical request through the boundary:
- The application sends an internal request: tenant, caller service, model or task class, messages or input, timeout budget, trace ID.
- The gateway authenticates the caller (service identity), not the end user — end-user auth usually already happened at the API gateway. Tenant identity must come from authenticated/trusted context or be validated against the caller’s authorization; do not treat an arbitrary caller-supplied tenant ID as trusted.
- Policy runs: is this tenant allowed this model? Is the budget exhausted? Is the request over a size cap?
- The adapter maps the internal request onto a provider call (and may select a provider if routing is enabled).
- The call is executed with timeouts. Retries, if any, follow a named policy.
- The response is passed back, with enough normalization to be usable, without stripping provider-specific fields the application needs.
- Telemetry and usage are recorded against tenant, service, model ID, and provider.
That is the whole job. If a step is not about reaching providers under policy, it probably belongs elsewhere.
Architecture
In a production AI system the gateway is one capability on the production AI stack, not the stack itself. AI System Architecture places it next to orchestration: the orchestrator decides what to ask a model; the gateway decides how that ask is sent.
Diagram: API gateway and AI gateway as different boundaries
flowchart TB
Client[Client]
APIGW[API Gateway]
Svc[Application / AI service]
AIGW[AI Gateway]
Models[Model providers]
Client --> APIGW
APIGW --> Svc
Svc --> AIGW
AIGW --> Models
The API gateway fronts your services. The AI gateway fronts model providers. They can share a process; they are different jobs.
The application remains the place for product behavior: prompts, retrieval, tools, and user-facing authorization. The AI gateway should not need to understand a support ticket or a document corpus.
Core responsibilities
Treat the list below as a menu, not a bill of materials. “Common” means teams often want it once a gateway exists. “Mandatory” means the concept stops being a gateway without it — almost nothing is mandatory.
| Responsibility | What it does | How common |
|---|---|---|
| Provider adapters | Map one internal request shape onto provider APIs | Common with two backends; not mandatory (one SDK can be the adapter) |
| Centralized credentials | Hold and rotate provider keys away from product repos | Common in multi-service orgs; a secrets manager may be enough |
| Caller authn/authz | Identify the service (and maybe tenant) allowed to call models | Needed if the gateway is a network hop beyond one app |
| Quotas and spend controls | Cap tokens, requests, or dollars per tenant/service | Common in shared platforms; not mandatory |
| Rate limiting | Limit request/token traffic before provider quotas are exceeded | Common at meaningful QPS; not mandatory |
| Model/provider policy | Allow-lists, region pins, “this tenant cannot use model X” | Common in multi-tenant systems; not mandatory |
| Request normalization | Stable internal schema (messages, tools, stream flag) | Common with multiple providers; not mandatory |
| Response normalization | Enough structure to consume results uniformly | Use carefully — over-normalizing hides provider semantics |
| Retries and fallbacks | Retry transient errors; optionally try another model/provider | Common, but retries have cost side effects |
| Timeouts | Bound every provider call | Strongly recommended whenever you call a network |
| Observability | Spans for provider, model ID, latency, status, tokens | Common; you will regret skipping it |
| Usage and cost attribution | Attach tokens/$ to tenant, feature, or route | Common when finance asks; not mandatory |
| Safety/security controls | PII redaction, block lists, egress controls at this boundary | Situational — not a substitute for AI Security |
| Tenant isolation | Prevent cross-tenant use of keys, quotas, and traces | Common in SaaS; skip for a single-tenant app |
Decision Trade-off
A gateway that only forwards HTTP with a stored API key is a credentials proxy. That can still be worth it. Do not pretend it is also routing, eval, and a safety platform until those jobs have owners and tests.
Timeouts are the closest thing to “always.” An unbounded provider call will eventually take the application down with it. Latency Optimization is how you budget the hop; the gateway is one place to enforce the budget.
Retries are not free. A retried completion usually bills again. Hedged requests (send two providers, take the first) can double cost on the happy path. Fallback is not equivalent behavior: a cheaper model is a different system, not a hot spare with the same answers. Pair fallbacks with evaluation if quality matters.
Observability belongs here because this is where provider latency and token counts are known. It does not replace application tracing. Propagate the same trace_id the API gateway and orchestrator already use — see Observability. Prompts, completions, tool arguments, and retrieved content may contain sensitive or tenant-isolated data. Record them only under an explicit retention, redaction, and access policy.
Spend controls can reduce unbounded spend. They do not automatically reduce unit cost. Cost Optimization is mostly routing, caching, prompt size, and model choice. A gateway can enforce a cap; it cannot make a frontier model cheap.
What Should NOT Live in the AI Gateway?
This is the section that prevents the gateway from becoming an accidental AI platform.
Application business logic. Pricing rules, ticket workflows, “is this user allowed to refund an order,” and domain prompts are product code. The gateway may receive a tenant ID; it should not interpret the ticket.
Agent orchestration. Tool loops, planning, memory writes, and “what should we do next” belong in the application or an orchestrator — see AI Agents and Tool Calling. The gateway may execute a model call that happens to include tool schemas. It should not own the loop.
RAG retrieval. Chunking, indexing, hybrid search, and ACL filters belong on the retrieval path — see RAG and Enterprise RAG Architecture. Stuffing retrieval inside the gateway hides authorization mistakes behind a “model access” service.
Document processing. Parsing, OCR, and ingestion pipelines are data-plane jobs. They may call the gateway for an extraction model. They should not run inside it.
Evaluation logic. Golden sets, graders, and offline quality gates are how you know a model or prompt change was not a regression — see Evaluation. Sampling traces from the gateway into an eval set is fine. Scoring answers in the gateway on the hot path is usually the wrong coupling.
Long-term application memory. Session stores, user profiles, and agent memory are application state — see Agent Memory. The gateway’s job is a call, not a corpus.
Training and fine-tuning. Job orchestration, datasets, and training clusters are a different discipline. A gateway might call a fine-tuned deployment; it does not train it.
All application authorization. End-user authn/authz, object-level ACLs, and “this document is visible to this role” are application and data-plane concerns. The gateway can enforce model-access policy (which model, whose provider key, which spend envelope). It cannot be the only place you check whether a user may see a record. AI Security starts before the model is called.
Warning
If the gateway repository contains your product prompts, retrievers, and agent graphs, you do not have a gateway. You have a monolith with a fashionable name. Platform teams should own provider access; product teams should own product behavior.
Step-by-Step Flow
Diagram: A completion request through an AI gateway
sequenceDiagram
participant App as AI service
participant GW as AI Gateway
participant Pol as Policy
participant Prov as Provider
App->>GW: complete(tenant, model, input, trace)
GW->>Pol: auth caller, quota, allow-list
Pol-->>GW: allow or deny
GW->>Prov: adapted request
Prov-->>GW: response or error
GW->>GW: record usage and span
GW-->>App: result plus model id
Policy runs before the provider call. Traces are written even when the call fails.
Walkthrough for a multi-service setup:
- API gateway authenticates the user, rate-limits the client, attaches
trace_idand tenant. That is ordinary HTTP work. - Application builds the prompt, optionally retrieves evidence, and decides it needs a model completion. It does not hold the provider API key.
- AI gateway checks that this service may spend against this tenant’s envelope, that the requested model is allowed, and that a timeout is set.
- Adapter sends the provider-native request. On classified transients such as
429/503, a production gateway may retry under a named, bounded policy (with exponential backoff/jitter where appropriate) or, if configured, try a fallback model. Honor providerRetry-Afterwhen supplied, and keep retries inside the caller’s remaining deadline. It records retry count. A fallback is a behavior change, not an equivalent spare. - Response returns to the application with
model_id, token counts, and the provider request ID. The application continues with its own guardrails and UX. - Async usage export feeds billing and cost dashboards. Failures still emit spans.
The application never needed to know which HTTP header the provider wanted. It did need the actual model_id that ran — hiding that breaks eval and incidents.
Real Production Example
Illustrative implementation. The following is a thin in-process gateway: one internal request type, named model-access policy, in-memory quota accounting, and a single retry on classified transients. It is not a product and is not production-ready. A shared proxy would expose the same complete() over HTTP.
This is illustrative pseudocode showing the architectural boundary. Provider-specific request formats, real timeout enforcement, retry backoff, distributed quota accounting, tracing SDK integration, persistence, and streaming are omitted for clarity.
from __future__ import annotations
import time
from dataclasses import dataclass
from typing import Protocol
class ProviderError(Exception):
def __init__(self, status: int, message: str):
super().__init__(message)
self.status = status
class Provider(Protocol):
name: str
async def complete(self, req: "CompletionRequest") -> "CompletionResult": ...
@dataclass(frozen=True)
class CompletionRequest:
tenant: str
caller: str
model: str
input: str
timeout_s: float # request contract; enforcement omitted below
trace_id: str
@dataclass(frozen=True)
class CompletionResult:
text: str
provider: str
model: str
input_tokens: int
output_tokens: int
class QuotaExceeded(Exception):
pass
class PolicyDenied(Exception):
pass
class AiGateway:
"""Provider-access boundary: policy, one retry, usage, no product logic."""
def __init__(
self,
providers: dict[str, Provider],
*,
allow: dict[str, set[str]],
daily_token_cap: dict[str, int],
usage: dict[tuple[str, str], int],
):
self.providers = providers
self.allow = allow
self.daily_token_cap = daily_token_cap
self.usage = usage
def _provider_for(self, model: str) -> Provider:
# Routing can replace this map. The gateway only needs a resolution step.
if model.startswith("local-"):
return self.providers["local"]
if model.startswith("claude"):
return self.providers["anthropic"]
return self.providers["openai"]
def _enforce(self, req: CompletionRequest) -> None:
allowed = self.allow.get(req.tenant, set())
if allowed and req.model not in allowed:
raise PolicyDenied(f"{req.tenant} cannot use {req.model}")
used = self.usage.get((req.tenant, _day()), 0)
if used >= self.daily_token_cap.get(req.tenant, 10_000_000):
raise QuotaExceeded(req.tenant)
async def complete(self, req: CompletionRequest) -> CompletionResult:
self._enforce(req)
provider = self._provider_for(req.model)
started = time.monotonic()
try:
result = await self._call(provider, req)
except ProviderError as exc:
self._emit(req, provider.name, error=str(exc), latency_s=time.monotonic() - started)
raise
tokens = result.input_tokens + result.output_tokens
key = (req.tenant, _day())
self.usage[key] = self.usage.get(key, 0) + tokens
self._emit(req, result.provider, tokens=tokens, latency_s=time.monotonic() - started)
return result
async def _call(self, provider: Provider, req: CompletionRequest) -> CompletionResult:
try:
return await provider.complete(req)
except ProviderError as exc:
if exc.status not in {429, 503}:
raise
# One immediate retry on classified transients. Backoff/jitter omitted.
return await provider.complete(req)
def _emit(self, req: CompletionRequest, provider: str, **fields) -> None:
print(
{
"trace_id": req.trace_id,
"tenant": req.tenant,
"caller": req.caller,
"model": req.model,
"provider": provider,
**fields,
}
)
def _day() -> str:
return time.strftime("%Y-%m-%d")
Real shared quotas require atomic/distributed accounting and a defined policy for estimated versus final token usage; this example uses in-memory accounting only to show the control flow.
What matters architecturally:
- Policy is named (
allow,daily_token_cap) and checked before the provider call. timeout_sis on the request contract. The example does not enforce it. Production gateways must bound every provider call.- The application still owns the prompt. The gateway never fetches documents or runs tools.
- Routing is a function (
_provider_for). Replacing it with a richer policy does not change the boundary. - One immediate retry is shown for classified
429/503errors. Production policy is bounded, retries only explicitly classified transients, uses exponential backoff/jitter where appropriate, and records retry count in telemetry. This example omits backoff and retry-count emission. - There is no hidden hedge that doubles spend on the happy path.
- Traces keep
modelandprovider. An abstraction that swallows those fields is how incidents go blind.
Extract this to a service when multiple codebases need the same policy. Until then, a module is a gateway.
Simple Architectures
These shapes are progressively more realistic. Stop at the first one that matches the actual organization.
A. Single application + single provider
One API, one model ID, keys in a secrets manager, timeouts and traces in the application.
A gateway is usually unnecessary. A provider SDK is the adapter. Adding a proxy adds latency and another component to page without removing duplication — there is nothing to duplicate yet.
B. Multiple services + multiple providers
Chat, batch jobs, and an eval runner talk to two hosted APIs and one self-hosted model. Each service previously shipped its own client.
A gateway starts to pay for itself as shared credentials, a stable internal request type, and one place to emit token counts. Routing can still be a static map (task → model_id).
C. Multi-tenant production system
Several products, billed tenants, and a finance requirement to attribute spend. Some tenants are pinned to a region or banned from a model tier.
Quotas, caller identity, and policy belong on the boundary. Tenant isolation here means model-access isolation (keys, envelopes, allow-lists), not document ACLs. Document ACLs stay in retrieval and application code.
D. Gateway + model routing
The gateway remains the control/integration boundary. A routing policy chooses which model or provider should handle this request — by task, latency class, cost envelope, or availability.
Diagram: Routing as a decision behind the gateway
flowchart TB
App[Application]
subgraph GW[AI Gateway]
Pol[Policy / auth / quotas]
Tel[Telemetry]
Retry[Retries]
Route[Model routing]
end
App --> GW
Route --> A[Provider A]
Route --> B[Provider B]
Route --> C[Provider C]
The gateway owns the boundary. Routing is one policy it may run — not a synonym for the gateway.
See Model Routing for the decision process itself. Treat routing as a versioned policy next to Cost Optimization and Latency Optimization, logged on every request.
Design Decisions
| Decision | Lean | Heavier |
|---|---|---|
| Where it runs | In-process client (one repo/team) | Shared proxy (many languages or teams) |
| How much to abstract | Thin wrappers when you need native features | One generic schema when callers must not know providers |
| Credentials | App role + secrets manager | Gateway-held keys when product repos must not hold them |
| Routing | Pinned model_id |
Policy table when task/cost/availability actually differ |
| Retries | Timeout only | One retry on 429/503 when provider blips hit the SLO |
| Fallback | Fail the request | Secondary model when availability beats identical answers |
| Spend caps | Invoice alerts | Hard per-tenant quotas when abuse is likely |
Engineering Insight
The lean column is production-grade for a surprising number of products. Heaviness is a response to organizational fan-out, not a maturity badge.
Comparisons
AI Gateway vs API Gateway
An API gateway is a general service/API boundary: terminate TLS, authenticate clients, rate-limit HTTP, route to application services, maybe inject identity headers.
An AI gateway is a model-provider boundary: provider credentials, token and model usage, provider-specific limits, request/response adaptation, and AI-specific cost, telemetry, and policy.
They overlap in boring ways (auth, rate limits, observability). Overlap does not make them the same object. An organization may run both. An AI gateway does not replace an API gateway; an API gateway does not know what a token budget is unless you teach it.
| Concern | API gateway | AI gateway |
|---|---|---|
| Sits in front of | Your application services | Model providers |
| Primary callers | End users, other products, browsers | Internal AI services |
| Auth | User/session/API keys for your APIs | Service identity to spend on models |
| Rate limiting | HTTP request rates | Tokens, RPM, provider quotas |
| Routing target | Which service handles the URL | Which model/provider handles the completion |
| Payload semantics | Opaque HTTP for most APIs | Messages, tools, tokens, streams |
| Cost unit | Usually infra, not per-token | Tokens and provider invoices |
| Failure vocabulary | 4xx/5xx of your API | Provider 429/503, context overflow, safety blocks |
You can implement AI-specific checks inside an API gateway plugin. You still have two jobs. Naming them separately keeps the public API from accumulating provider SDKs, and keeps the model boundary from accumulating user-login flows.
AI Gateway vs Model Routing
AI gateway = the control and integration boundary (how we call providers, under what policy, with what telemetry).
Model routing = the decision mechanism for which model or provider should handle a request.
A gateway may contain or invoke routing. The concepts are not identical. A single-provider app can route between model tiers inside application code with no gateway. A gateway can exist solely for credentials, timeouts, and traces with one model behind it.
| AI gateway | Model routing | |
|---|---|---|
| Question | How do we talk to providers uniformly and safely? | Which model/provider should handle this request? |
| Kind of thing | Boundary / adapter / policy enforcement point | Decision policy |
| Can exist alone? | Yes — one model, shared keys and traces | Yes — pinned IDs or if/else in the app |
| Needs eval? | Helpful for fallbacks | Yes, if routes differ in quality |
| Failure if confused | “We bought a gateway, so we must be routing” | “We route in five services with no shared adapter” |
Remember
If you cannot say which model ID served a request, you are not routing — you are hoping. If you cannot say which component holds provider keys, you do not yet have a gateway; you have copies.
Common Mistakes
-
Creating a gateway before there is a multi-service or multi-provider problem. You add an outage domain and a deploy for a key that lived in one environment file.
-
Turning the gateway into a monolith. Prompts, retrievers, agents, and billing UI accumulate because “it is the AI service.” Split along jobs, not along a buzzword.
-
Hiding provider semantics behind an overly generic abstraction. Tool calling, vision inputs, logprobs, and structured outputs are not identical across providers. A lowest-common-denominator schema will either drop fields you need or pretend they are portable when they are not.
-
Centralizing business authorization in the gateway. “May this user refund order 123?” is not a model-access question. Put object ACLs next to the data. The gateway may only know “tenant T may spend on model M.”
-
Retries that duplicate expensive requests. Completions are often not idempotent for cost. Cap retries, log them, and never silently hedge.
-
Treating fallback as equivalent behavior. A backup model is a different product behavior. Tell the application (and eval) which model actually ran.
-
Losing observability in the abstraction. If spans do not include provider, model ID, token counts, and retry count, the gateway is a blind fold, not a boundary. See Observability.
-
Buying a gateway to postpone pinning a model ID. A proxy without a named route and a tiny golden set is another hop. See Production AI Stack.
-
Assuming the gateway makes switching providers easy. It can make transport switching easier. Prompts, tools, tokenization, and quality still change. You still need eval.
Where It Breaks Down
Streaming and multimodal I/O. Byte streams, audio sessions, and image batches may not fit a single “chat completions” adapter. Keep the boundary, but do not force every modality through a text-shaped schema.
Provider-unique capabilities. If the product is a provider feature (a specific computer-use API, a unique cache prefix, a proprietary safety classifier), a generic gateway will lag. Call that provider more natively and isolate the exception.
Ultra-low latency paths. An extra hop plus policy checks can dominate a 50ms classifier. In-process is still a gateway. A remote proxy may not belong on that path.
Research and one-off jobs. Notebooks and paper reproductions should call a provider. Do not require a platform ticket to try a model ID.
Organizational bottlenecks. A platform team that must approve every prompt because the gateway owns prompts will stall products. Own keys and quotas centrally; own prompts locally.
When to Use an AI Gateway
Use one when several of these are true at once — not when a vendor diagram includes the box.
- Multiple services call models and are already copying clients, retries, or key handling
- Multiple model providers (hosted and/or self-hosted) must look like one internal contract
- Centralized provider credentials are desirable so product repos do not hold long-lived provider keys
- Spend and usage must be attributed or capped per tenant, team, or feature
- Provider switching or failover is an operational requirement you can actually test
- Consistent telemetry (model ID, tokens, provider latency) is missing because each service emits something different
- Tenant quotas or model allow-lists must be enforced in one place rather than by convention
A useful test: if we delete this component, which copied logic would immediately reappear in three repositories? If the answer is “nothing,” you do not need it yet.
Decision tree: is an AI gateway earned?
flowchart TD
Start[Who calls models?] --> Multi{Many services or providers?}
Multi -->|No| Direct[SDK + timeouts + traces]
Multi -->|Yes| Keys{Shared keys or spend policy?}
Keys -->|No| Maybe[Shared client may suffice]
Keys -->|Yes| Gw[Gateway pays off]
Gw --> Route{Task quality or cost differ?}
Route -->|Yes| RoutePol[Add routing policy]
Route -->|No| Thin[Keep a thin adapter]
Fan-out of services, providers, or policy is the signal. Routing is optional on top.
When NOT to Use an AI Gateway
A gateway adds an operational boundary: another deploy, another timeout, another way to fail before you even reach the model.
Do not introduce one when:
- You have one model provider, one service, simple credentials, and low operational complexity
- Direct provider integration with timeouts, a secrets manager, and traces would be simpler — because it is simpler
- You cannot name the duplication the gateway would remove
- You are using the gateway as a way to avoid pinning model IDs or writing a ten-case eval set
- You plan to put RAG, agents, or product workflows “in the platform” so application teams do not have to think
This guide should actively discourage unnecessary architecture. A traced single-provider API can be more production-ready than an unmeasured multi-provider mesh.
Warning
“We might use another provider later” is not a reason to add a hop today. A stable internal function around the SDK is enough to keep that door open.
Running in Production
Best Practice
Pin model IDs. Bound every provider call. Attribute tokens to tenant and caller. Log retries. Keep product logic out of the gateway repository.
| Dimension | Guidance |
|---|---|
| Scaling | Size capacity around concurrent provider work and downstream limits, not just end-user request rate. Do not couple it to retrieval QPS. |
| Latency | Budget the hop (connection reuse, region). Do not put RAG or tool loops on this path. Stream through when the application streams. When proxying a stream, propagate cancellation/disconnect and preserve partial-output semantics; do not buffer the entire response merely to normalize it. |
| Cost | Caps prevent runaway spend; they do not optimize unit cost. Record tokens before and after retries. |
| Monitoring | Spans: trace_id, caller, tenant, provider, model_id, status, tokens, retry count, fallback used. |
| Evaluation | Fallback and routing changes need golden-set coverage. The gateway should emit the fields eval needs, not run eval itself. |
| Security | Provider keys in a vault; service-to-service auth to the gateway; no end-user secrets in prompts. See AI Security. |
| Change | Canary a model ID or provider adapter. Rollback independently of application deploys. |
Production checklist
- Provider credentials not stored in product application repos
- Timeouts on every provider call
-
model_idand provider name on every trace - Token/cost attribution per tenant and caller
- Retry policy named, capped, and logged; honor
Retry-After; stay within the caller deadline - Fallbacks treated as behavior changes, not silent equivalents
- Gateway does not own prompts, retrieval, or agent loops
- Streaming proxies propagate cancellation/disconnect and partial output; do not buffer the full response merely to normalize it
- Load and error SLOs for the gateway itself (it can take the product down)
Interview Questions
-
Why would an application put an AI gateway in front of model providers?
To stop copying provider integrations, credentials, policy, and telemetry across services — not because production requires a proxy. -
When is a gateway unnecessary?
One service, one provider, simple secrets, low complexity: call the SDK with timeouts and traces. -
How is an AI gateway different from an API gateway?
API gateways front your services (HTTP, user auth, URL routing). AI gateways front model providers (keys, tokens, adapters, model policy). Both may exist. -
How is a gateway different from model routing?
The gateway is the integration boundary. Routing is the policy that chooses a model. Either can exist without the other. -
What should generally stay out of the gateway?
Product business rules, document retrieval/RAG, agent orchestration, training workflows, and object-level user authorization. Exceptions can exist when a capability is explicitly part of provider-access policy, but the gateway should not become the application's general AI platform. -
Why can retries be a production incident?
Many model calls are billable per attempt. Unbounded retries amplify cost and can duplicate side effects if tools are in the same path (they should not be). -
Does a gateway make provider switching easy?
It can stabilize transport. It does not make models interchangeable. Quality, tools, and tokenization still need eval. -
In-process or a service?
Stay in-process until multiple codebases or languages need the same policy. A module can be a gateway.
Related Guides
This guide is the provider-access boundary. Production AI Stack is the capability map. AI System Architecture is the platform blueprint.
Architecture:
- Production AI Stack — where the gateway sits among other capabilities
- AI System Architecture — layered platform, orchestration, contracts
- Model Routing — which eligible model or provider should handle a request
- AI Copilot Architecture — application pattern that may call the gateway
- Enterprise RAG Architecture — retrieval stays off this boundary
Foundations:
Operations:
- Cost Optimization · Latency Optimization · Observability · AI Security · Evaluation · Guardrails · Caching
Agents (keep them out of the gateway):
Rankings: Best AI APIs — hosted access, not a mandate to add a gateway
Learning path: Become an AI Engineer
Diagram: Recommended reading around this guide
flowchart LR
LLM[LLM Concepts] --> Stack[Production AI Stack]
Stack --> GW[AI Gateway]
GW --> Route[Model Routing]
GW --> Arch[System Architecture]
GW --> Ops[Cost and Obs]
Read the stack map first if you need composition context; this page is only the provider boundary.
Learning Path
Prerequisites: Large Language Models
Next topics: Production AI Stack · AI System Architecture · Model Routing · AI Copilot Architecture · Cost Optimization · Observability
Estimated time: 35 min · Difficulty: Advanced
FAQs
Is an AI gateway the same as an API gateway?
No. An API gateway is a general HTTP/service boundary in front of your applications. An AI gateway is a model-provider boundary: adapters, provider credentials, token usage, and model policy. They can overlap in implementation and you may run both.
Do I need an AI gateway for a single-model application?
Usually not. One service and one provider can use the vendor SDK with timeouts, secrets, and traces. Extract a gateway when that integration is being copied or when shared policy appears.
Does an AI gateway perform model routing?
It might. Routing is a separate decision: which model should handle the request. A gateway is a place that decision can run. A gateway without routing is still a gateway; routing without a gateway is still routing.
Does an AI gateway replace an SDK?
No. Something still has to speak HTTP to providers. That may be a provider SDK inside the gateway’s adapters, or a generated client in front of your proxy. Application teams should depend on your stable client, not on five vendor SDKs — unless they only have one vendor.
Should RAG run inside the gateway?
No. Retrieval, ACLs, and indexing are a data plane. The gateway may complete a prompt that already contains retrieved evidence. See RAG.
Should agent orchestration run inside the gateway?
No. Loops, tools, and memory are application control flow. The gateway executes model calls those loops need. See AI Agents.
Can an AI gateway reduce model costs?
It can cap spend and make cost visible. Lower unit cost comes from smaller prompts, cheaper models, caching, and routing policy — see Cost Optimization. A gateway does not make tokens cheaper by existing.
Does an AI gateway make switching providers easy?
It can reduce the cost of changing transport (one adapter instead of five SDKs). It does not make two models interchangeable. Expect prompt, tool, and quality work, plus eval, whenever the backend changes.
Does an AI gateway automatically improve security?
No. It can centralize provider keys and enforce model-access policy. Application authorization, prompt injection defenses, and data ACLs remain application and retrieval problems — see AI Security.
Can the gateway be in-process?
Yes. A module with adapters, policy, and traces is an AI gateway. Promote it to a service when multiple codebases need the same boundary.
Will a gateway hide provider outages?
Only if you add failover and accept that fallback models behave differently. Timeouts stop hung calls; they do not create capacity.
References
- OpenAI API Documentation
- Anthropic API Documentation
- Google AI for Developers
- OpenTelemetry Documentation
- LiteLLM Documentation
Further Reading
- Production AI Stack
- AI System Architecture
- Model Routing
- Cost Optimization
- Observability
- Best AI APIs
Key Takeaways
- An AI gateway is a centralized runtime boundary for accessing and governing model providers — adapter plus policy, not a mandatory product.
- It exists to stop duplicated provider integrations (keys, quotas, telemetry, retries), not to host the rest of the AI stack.
- API gateway ≠ AI gateway. One fronts your services; the other fronts model providers. You may need both.
- Gateway ≠ routing. Routing is a decision the gateway may execute.
- RAG, agents, eval, memory, and business authorization do not belong here.
- One service and one provider should usually call the SDK directly.
- Do not claim a gateway guarantees lower cost or automatic security — it can enforce caps and centralize keys; that is all.