vLLM
FreePopularHigh-throughput open LLM inference engine.
High-throughput LLM inference engine with PagedAttention.
Why vLLM matters
vLLM is the default open serving stack for many self-hosted LLM deployments. PagedAttention and continuous batching made high-throughput GPU inference practical—so model choice and serving choice are now separate decisions. vLLM 0.30.0 (2026-09-22) adds DeepSeek-V4.1-Flash serving; 0.31.0 (2026-10-05) adds `vllm preload` fast restarts and SM100 support for that model, with breaking changes to fp8 quantization naming and per-request multimodal kwargs. Model Runner V2 remains the default from 0.29.0.
Last reviewed: 7 October 2026
When to choose vLLM
Decision guidance for architects—not a feature list.
Best for
- Self-hosted LLM serving
- High-throughput inference
- Open-weight model deployment
- GPU utilization optimization
Avoid if
- You only need a hosted API (OpenAI/Anthropic/etc.)
- You want a zero-ops consumer chat product, not a serving engine
Strengths
Qualitative snapshot for architects—not a public ranking.
- Throughput★★★★★
- Open-weight serving★★★★★
- GPU efficiency★★★★★
- Managed SaaS★☆☆☆☆
- Beginner setup★★★☆☆
Ecosystem
Competes with
- Ollama
Ollama wins on local DX; vLLM wins on production throughput and batching.
Alternative to
- Transformers
Transformers for training/dev; vLLM when you need high-QPS serving.
Works with
- Ollama
Ollama for local DX; vLLM for production throughput (different tiers).
Recommended for
- Latency Optimization
Serving knobs and batching dominate real-world LLM latency.
Often paired with
- Meta
Llama open weights are commonly served with vLLM.
- NVIDIA
vLLM targets high-throughput serving on NVIDIA GPUs.
- Hugging Face
Hub models are frequently deployed through vLLM.
How vLLM evolved
Key moments in chronological order.
- Release
Adds vllm preload for fast restarts and DeepSeek-V4.1-Flash on SM100. Breaking: per-request mm_processor_kwargs/media_io_kwargs require --trust-request-mm-kwargs, tokenizer_mode="slow" is removed, and quantization="fp8" becomes fp8_per_tensor.
- Release
Adds DeepSeek-V4.1-Flash, Fast Start GPU weight cache via --load-format ipc_cache, and HiSparse host-tier KV for sparse MLA. Scale-out endpoints on vllm serve now require --enable-scale-out.
- Release
Model Runner V2 becomes the default for all models (MRV1 remains for some ROCm paths). Adds Hy4-preview, Qwen3.8-Flash-Next, and Kimi K3 NVFP4; DeepSeek V4 serving improvements. Breaking: deprecated architectures removed; prefer vllm serve.
- Platform
Official vLLM TT Plugin serves Llama, Qwen, Mistral, Gemma, DeepSeek V3, and GPT-OSS on Tenstorrent accelerators via the out-of-tree platform plugin. The OpenAI-compatible serving API is unchanged; model code lives in TT-Metal.
- Release
Muse Glimmer support plus KV-cache offload and speculative-decoding improvements, with continued Qwen and ROCm coverage.
- Release
Day-0 Kimi K3 stack plus Qwen3.5/3.8-class model support; PyTorch 2.13 upgrade.
- Platform
Production hardening
Metrics, continuous batching, and model coverage deepen for real traffic.
- Product
LoRA / multi-adapter serving
Serving many fine-tuned adapters efficiently becomes a production requirement.
- Platform
Ecosystem default for open serving
Becomes a standard choice next to TGI and TensorRT-LLM in many stacks.
- API
OpenAI-compatible serving
Drop-in API compatibility accelerates migration from hosted APIs to self-host.
- Open source
PagedAttention / vLLM rise
Open serving engine makes high-throughput LLM inference widely accessible.
Tool Info
Overview
vLLM is a serving engine optimized for high-throughput LLM inference.
Provides OpenAI-compatible APIs for production deployments.
vLLM 0.28.0 (2026-08-26) adds Muse Glimmer support plus KV-cache offload and speculative-decoding improvements.
vLLM TT Plugin (2026-09-07) adds Tenstorrent accelerators through the out-of-tree platform plugin; serving API is unchanged.
vLLM 0.29.0 (2026-09-09) makes Model Runner V2 the default for all models; prefer vllm serve.
vLLM 0.30.0 (2026-09-22) adds DeepSeek-V4.1-Flash, Fast Start GPU weight cache (`--load-format ipc_cache`), and HiSparse. Scale-out endpoints on vllm serve require `--enable-scale-out`.
vLLM 0.31.0 (2026-10-05) adds `vllm preload` for fast restarts and DeepSeek-V4.1-Flash on SM100 (Blackwell). Breaking changes: per-request `mm_processor_kwargs`/`media_io_kwargs` now require `--trust-request-mm-kwargs`, `tokenizer_mode="slow"` is removed, and `quantization="fp8"` becomes `fp8_per_tensor`.
Pricing
Pros
- Best-in-class throughput
- PagedAttention
- Wide model support
Best For
When NOT to Use
- GPU required
- Ops complexity
Community Insights
Real implementation experiences shared by AI practitioners.
Loading practitioner experiences…
Related Tools
Alternatives
Tags
Related Guides
- Fine-Tuning
Engineer LLM fine-tuning with clear prompting and RAG boundaries, governed datasets, LoRA and QLoRA, evaluation gates, and production operations.
- Large Language Models
Learn how LLMs like GPT, Claude, and Llama process and generate human language at scale.
- Llama Models
Meta Llama open-weight models — Llama 4 Scout/Maverick for new deployments, Llama 3.x baselines, plus Muse Spark 1.3 (closed API / Muse Code) and Muse Glimmer (on-device open 30B).
- DeepSeek Models
DeepSeek models — V4.1-Flash (deepseek-flash) as the live multimodal Flash SKU, V4-Pro-0813 still on deepseek-v4-pro after 2026-09-14, V3/R1 open checkpoints, and peak/off-peak pricing from 2026-09-10.
- Sarvam Models
Sarvam AI models — Sarvam-M (24B), Sarvam 30B (2.4B active), and Sarvam 105B (10.3B active), Apache-2.0 weights trained for Indian languages.
Stay Updated
Get the latest AI news, tools, and engineering guides delivered to your inbox.
Subscribe to Newsletter