SarvamOpen SourceReasoningCoding
Sarvam 105B
Sarvam’s flagship open-weight reasoning model. A Mixture-of-Experts transformer with 10.3B active parameters, trained from scratch in India, and used to power the Indus assistant.
Tool calling · Thinking · Coding
Last reviewed: 25 September 2026
Overview
Sarvam’s flagship open-weight reasoning model. A Mixture-of-Experts transformer with 10.3B active parameters, trained from scratch in India, and used to power the Indus assistant.
Capabilities
- Vision: No
- Audio: No
- Tool calling: Yes
- Thinking: Yes
- MCP: No
- Coding: Yes
- Structured output: No
Technical specifications
- Provider
- Sarvam
- License
- Apache-2.0
- Context window
- 128K
- Parameters
- MoE 105B total / 10.3B active
- Architecture
- Mixture-of-Experts, 128 experts, top-8 routing, Multi-head Latent Attention
- Release
- 2026-03-06
- Modalities
- Text
- Vision
- No
- Audio
- No
- Tool calling
- Yes
- Thinking
- Yes
- MCP
- No
- Open weights
- Yes
- API
- Yes
- Pricing (input)
- Sarvam API (verify dashboard)
- Pricing (output)
- Sarvam API (verify dashboard)
Supported modalities
Text
Context window
128K (128,000 tokens)
Pricing
Input: Sarvam API (verify dashboard)
Output: Sarvam API (verify dashboard)
Open weights on Hugging Face and AI Kosh. API access is separate from the Apache-2.0 weights.
Availability
API: Yes
Chat UI: Yes
Open weights: Yes
Indus chat uses 105B. Self-host with Transformers, vLLM, or SGLang. Knowledge cutoff stated by Sarvam as June 2025.
Use cases
- Indian-language assistants
- Agentic tool-use workflows
- Math and coding reasoning
- Open-weight sovereign deployments
Strengths
- Apache-2.0 weights trained in India
- Strong published reasoning and agentic results for its class
- Indic tokenizer covering 22 scheduled languages
Limitations
- Text-only; vision and speech are separate Sarvam products
- Smaller global tooling ecosystem than Llama or DeepSeek
Related guides
Related benchmarks
Related research
- Open-sourcing Sarvam 30B and 105B
Original Paper
Related GitHub
Related tools
Related rankings
Companies
Explore more models
- Sarvam 30BSarvamSarvam’s efficient open-weight reasoning MoE. About 2.4B active parameters, built for real-time conversational agents on the Samvaad platform and for local or GPU-constrained deployment.
- Sarvam-MSarvamSarvam’s earlier multilingual model: a 24B hybrid-reasoning post-train of Mistral Small, released under Apache-2.0 in May 2025. API id sarvam-m.
- DeepSeek R1DeepSeekDeepSeek’s reasoning-focused model trained with reinforcement learning for multi-step math, science, and coding problem solving.
- DeepSeek V3DeepSeekDeepSeek’s MoE general model — strong open-weight performance on coding and knowledge tasks with competitive API pricing.
- DeepSeek V4DeepSeekDeepSeek’s V4 generation — live Flash is V4.1-Flash (API id deepseek-flash): Causal Encoder–Decoder MoE, 552B backbone (8B active prefill / 16B decode), native image+text, and 1M context. MIT weights on Hugging Face. Legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp aliases temporarily route here. deepseek-v4-pro still serves V4-Pro-0813 after 2026-09-14 at unchanged Pro rates (DeepSeek withdrew the planned Flash reroute).
- Kimi K3Moonshot AIMoonshot’s Kimi K3 — 2.8T MoE (104B active) open-weight multimodal agentic model with 1M context, native vision, and strong long-horizon coding. Weights on Hugging Face under the Kimi K3 License.