vLLM

FreePopular

High-throughput open LLM inference engine.

High-throughput LLM inference engine with PagedAttention.

Why vLLM matters

vLLM is the default open serving stack for many self-hosted LLM deployments. PagedAttention and continuous batching made high-throughput GPU inference practical—so model choice and serving choice are now separate decisions. vLLM 0.30.0 (2026-09-22) adds DeepSeek-V4.1-Flash serving; 0.31.0 (2026-10-05) adds `vllm preload` fast restarts and SM100 support for that model, with breaking changes to fp8 quantization naming and per-request multimodal kwargs. Model Runner V2 remains the default from 0.29.0.

Open SourceAPISelf-hostedPython SDK

Last reviewed: 7 October 2026

When to choose vLLM

Decision guidance for architects—not a feature list.

Best for

  • Self-hosted LLM serving
  • High-throughput inference
  • Open-weight model deployment
  • GPU utilization optimization

Avoid if

  • You only need a hosted API (OpenAI/Anthropic/etc.)
  • You want a zero-ops consumer chat product, not a serving engine

Strengths

Qualitative snapshot for architects—not a public ranking.

  • Throughput★★★★★
  • Open-weight serving★★★★★
  • GPU efficiency★★★★★
  • Managed SaaS★☆☆☆☆
  • Beginner setup★★★☆☆

Ecosystem

Competes with

  • Ollama

    Ollama wins on local DX; vLLM wins on production throughput and batching.

Alternative to

  • Transformers

    Transformers for training/dev; vLLM when you need high-QPS serving.

Works with

  • Ollama

    Ollama for local DX; vLLM for production throughput (different tiers).

Recommended for

Often paired with

  • Meta

    Llama open weights are commonly served with vLLM.

  • NVIDIA

    vLLM targets high-throughput serving on NVIDIA GPUs.

  • Hugging Face

    Hub models are frequently deployed through vLLM.

How vLLM evolved

Key moments in chronological order.

  1. Release

    vLLM 0.31.0

    Adds vllm preload for fast restarts and DeepSeek-V4.1-Flash on SM100. Breaking: per-request mm_processor_kwargs/media_io_kwargs require --trust-request-mm-kwargs, tokenizer_mode="slow" is removed, and quantization="fp8" becomes fp8_per_tensor.

  2. Release

    vLLM 0.30.0

    Adds DeepSeek-V4.1-Flash, Fast Start GPU weight cache via --load-format ipc_cache, and HiSparse host-tier KV for sparse MLA. Scale-out endpoints on vllm serve now require --enable-scale-out.

  3. Release

    vLLM 0.29.0

    Model Runner V2 becomes the default for all models (MRV1 remains for some ROCm paths). Adds Hy4-preview, Qwen3.8-Flash-Next, and Kimi K3 NVFP4; DeepSeek V4 serving improvements. Breaking: deprecated architectures removed; prefer vllm serve.

  4. Platform

    vLLM Tenstorrent plugin

    Official vLLM TT Plugin serves Llama, Qwen, Mistral, Gemma, DeepSeek V3, and GPT-OSS on Tenstorrent accelerators via the out-of-tree platform plugin. The OpenAI-compatible serving API is unchanged; model code lives in TT-Metal.

  5. Release

    vLLM 0.28.0

    Muse Glimmer support plus KV-cache offload and speculative-decoding improvements, with continued Qwen and ROCm coverage.

  6. Release

    vLLM 0.27.0

    Day-0 Kimi K3 stack plus Qwen3.5/3.8-class model support; PyTorch 2.13 upgrade.

  7. Platform

    Production hardening

    Metrics, continuous batching, and model coverage deepen for real traffic.

  8. Product

    LoRA / multi-adapter serving

    Serving many fine-tuned adapters efficiently becomes a production requirement.

  9. Platform

    Ecosystem default for open serving

    Becomes a standard choice next to TGI and TensorRT-LLM in many stacks.

  10. API

    OpenAI-compatible serving

    Drop-in API compatibility accelerates migration from hosted APIs to self-host.

  11. Open source

    PagedAttention / vLLM rise

    Open serving engine makes high-throughput LLM inference widely accessible.

Tool Info

Categories
Serving · Infrastructure
Developer
vLLM
License
Open Source
Official Website

Overview

vLLM is a serving engine optimized for high-throughput LLM inference.

Provides OpenAI-compatible APIs for production deployments.

vLLM 0.28.0 (2026-08-26) adds Muse Glimmer support plus KV-cache offload and speculative-decoding improvements.

vLLM TT Plugin (2026-09-07) adds Tenstorrent accelerators through the out-of-tree platform plugin; serving API is unchanged.

vLLM 0.29.0 (2026-09-09) makes Model Runner V2 the default for all models; prefer vllm serve.

vLLM 0.30.0 (2026-09-22) adds DeepSeek-V4.1-Flash, Fast Start GPU weight cache (`--load-format ipc_cache`), and HiSparse. Scale-out endpoints on vllm serve require `--enable-scale-out`.

vLLM 0.31.0 (2026-10-05) adds `vllm preload` for fast restarts and DeepSeek-V4.1-Flash on SM100 (Blackwell). Breaking changes: per-request `mm_processor_kwargs`/`media_io_kwargs` now require `--trust-request-mm-kwargs`, `tokenizer_mode="slow"` is removed, and `quantization="fp8"` becomes `fp8_per_tensor`.

Pricing

Free tier available
Free (open source)
  • Best-in-class throughput
  • PagedAttention
  • Wide model support
Production LLM servingHigh-throughput inferenceOpenAI-compatible API
  • GPU required
  • Ops complexity

Real implementation experiences shared by AI practitioners.

Loading practitioner experiences…

Tags

#inference#serving#gpu

Related Guides

Stay Updated

Get the latest AI news, tools, and engineering guides delivered to your inbox.

Subscribe to Newsletter