llama.cpp
FreeHigh-performance C/C++ inference for LLaMA and GGUF models on CPU and GPU.
Tool Info
Overview
llama.cpp is the foundational inference engine for running quantized LLMs locally.
It powers Ollama, LM Studio, and many edge deployments.
Essential when you need maximum control over inference.
llama-server has a `/v1/systemone` endpoint for decision models (PR #29818): send a state plus typed questions and get a probability per option in one forward pass, with zero output tokens. It follows the System One format introduced with TypeSafe’s Jev and launched with official GGUFs for Laya, Julia-1, Lev, OpenJev, and Kev. Streaming is not supported. The endpoint first landed on master on 2026-10-02 and shipped in the stable v0.6.0 release on 2026-10-05.
v0.6.0 (2026-10-05) also adds Clef (text and vision) and Nimble model support and a new `llama_batch_ext` API, and bumps the session file format, so saved sessions from older builds must be regenerated.
Features
- GGUF format
- CPU and GPU backends
- Quantization support
- Server mode
- Decision-model endpoint (/v1/systemone)
Pricing
Pros
- Extremely efficient
- Runs on consumer hardware
- Massive community
Best For
When NOT to Use
- Lower-level than Ollama
- Manual model management
Typical Users
Community Insights
Real implementation experiences shared by AI practitioners.
Loading practitioner experiences…
Related Tools
Alternatives
Tags
Related Guides
- Large Language Models
Learn how LLMs like GPT, Claude, and Llama process and generate human language at scale.
- Decision Models vs. LLMs
Decision models answer typed questions (pick one, score, yes/no) with a probability for every option and no autoregressive decoding. How they differ from LLM structured outputs and classifiers, the /v1/systemone contract, OpenAI’s Decisions API (public beta), and production patterns for routing, triage, and agent gating.
Stay Updated
Get the latest AI news, tools, and engineering guides delivered to your inbox.
Subscribe to Newsletter