llama.cpp

Free

High-performance C/C++ inference for LLaMA and GGUF models on CPU and GPU.

Open SourceSelf-hostedC++ SDKPython SDK

Tool Info

Categories
Serving
Developer
ggml-org
License
Open Source

Overview

llama.cpp is the foundational inference engine for running quantized LLMs locally.

It powers Ollama, LM Studio, and many edge deployments.

Essential when you need maximum control over inference.

llama-server has a `/v1/systemone` endpoint for decision models (PR #29818): send a state plus typed questions and get a probability per option in one forward pass, with zero output tokens. It follows the System One format introduced with TypeSafe’s Jev and launched with official GGUFs for Laya, Julia-1, Lev, OpenJev, and Kev. Streaming is not supported. The endpoint first landed on master on 2026-10-02 and shipped in the stable v0.6.0 release on 2026-10-05.

v0.6.0 (2026-10-05) also adds Clef (text and vision) and Nimble model support and a new `llama_batch_ext` API, and bumps the session file format, so saved sessions from older builds must be regenerated.

Features

  • GGUF format
  • CPU and GPU backends
  • Quantization support
  • Server mode
  • Decision-model endpoint (/v1/systemone)

Pricing

Free tier available
Free (open source)
  • Extremely efficient
  • Runs on consumer hardware
  • Massive community
Local CPU inferenceGGUF model servingEdge deployment
  • Lower-level than Ollama
  • Manual model management
ML engineersEdge deployersPrivacy-focused teams

Real implementation experiences shared by AI practitioners.

Loading practitioner experiences…

Tags

#serving#inference#local

Related Guides

Stay Updated

Get the latest AI news, tools, and engineering guides delivered to your inbox.

Subscribe to Newsletter