SarvamOpen SourceReasoningCodingSmall Models
Sarvam 30B
Sarvam’s efficient open-weight reasoning MoE. About 2.4B active parameters, built for real-time conversational agents on the Samvaad platform and for local or GPU-constrained deployment.
Tool calling · Thinking · Coding
Last reviewed: 25 September 2026
Overview
Sarvam’s efficient open-weight reasoning MoE. About 2.4B active parameters, built for real-time conversational agents on the Samvaad platform and for local or GPU-constrained deployment.
Capabilities
- Vision: No
- Audio: No
- Tool calling: Yes
- Thinking: Yes
- MCP: No
- Coding: Yes
- Structured output: No
Technical specifications
- Provider
- Sarvam
- License
- Apache-2.0
- Parameters
- MoE 30B total / 2.4B active
- Architecture
- Mixture-of-Experts, 128 experts, top-6 routing, grouped-query attention
- Release
- 2026-03-06
- Modalities
- Text
- Vision
- No
- Audio
- No
- Tool calling
- Yes
- Thinking
- Yes
- MCP
- No
- Open weights
- Yes
- API
- Yes
- Pricing (input)
- Sarvam API (verify dashboard)
- Pricing (output)
- Sarvam API (verify dashboard)
Supported modalities
Text
Context window
Not specified
Pricing
Input: Sarvam API (verify dashboard)
Output: Sarvam API (verify dashboard)
Open weights on Hugging Face and AI Kosh. Sarvam reports higher throughput per GPU than a Qwen3 baseline on H100-class hardware.
Availability
API: Yes
Chat UI: Yes
Open weights: Yes
Powers Samvaad. Sarvam also publishes an MXFP4 path for Apple Silicon local inference.
Use cases
- Real-time Indian-language voice agents
- Cost-sensitive self-hosting
- Coding and math on a small active footprint
- Laptop-class local experiments
Strengths
- Low active compute for a reasoning MoE
- Apache-2.0 and Indic-efficient tokenizer
- Documented local and vLLM/SGLang serving paths
Limitations
- Trails 105B on the hardest agentic benchmarks Sarvam publishes
- Text model; telephony audio is a surrounding product stack
Related guides
Related benchmarks
Related research
- Open-sourcing Sarvam 30B and 105B
Original Paper
Related GitHub
Related tools
Related rankings
Companies
Explore more models
- Sarvam 105BSarvamSarvam’s flagship open-weight reasoning model. A Mixture-of-Experts transformer with 10.3B active parameters, trained from scratch in India, and used to power the Indus assistant.
- Sarvam-MSarvamSarvam’s earlier multilingual model: a 24B hybrid-reasoning post-train of Mistral Small, released under Apache-2.0 in May 2025. API id sarvam-m.
- Muse GlimmerMetaMeta Superintelligence Labs’ Muse Glimmer — Apache-2.0 ~30B dense multimodal agent model for on-device and single-GPU local agents. Sibling to closed Muse Spark; distinct from Llama 4.
- PhiMicrosoftMicrosoft’s Phi family of small language models — high capability per parameter for on-device, edge, and cost-sensitive deployments.
- DeepSeek R1DeepSeekDeepSeek’s reasoning-focused model trained with reinforcement learning for multi-step math, science, and coding problem solving.
- DeepSeek V3DeepSeekDeepSeek’s MoE general model — strong open-weight performance on coding and knowledge tasks with competitive API pricing.