MicrosoftOpen SourceSmall ModelsCodingReasoning
Phi
Microsoft’s Phi family of small language models — high capability per parameter for on-device, edge, and cost-sensitive deployments.
Vision · Tool calling · Coding
Last reviewed: 24 July 2026
Overview
Microsoft’s Phi family of small language models — high capability per parameter for on-device, edge, and cost-sensitive deployments.
Capabilities
- Vision: Yes
- Audio: No
- Tool calling: Yes
- Thinking: No
- MCP: No
- Coding: Yes
- Structured output: Yes
Technical specifications
- Provider
- Microsoft
- License
- MIT (varies by release)
- Context window
- 128K
- Parameters
- Family (3B–14B class)
- Architecture
- Dense Transformer (SLM)
- Release
- 2024–2025
- Modalities
- Text, Image
- Vision
- Yes
- Audio
- No
- Tool calling
- Yes
- Thinking
- No
- MCP
- No
- Open weights
- Yes
- API
- Yes
- Pricing (input)
- Self-host / Azure
- Pricing (output)
- Self-host / Azure
Supported modalities
Text · Image
Context window
128K (128,000 tokens)
Pricing
Input: Self-host / Azure
Output: Self-host / Azure
Small footprint reduces serving cost dramatically.
Availability
API: Yes
Chat UI: No
Open weights: Yes
Hugging Face + Azure AI model catalog.
Use cases
- On-device and edge assistants
- High-throughput classification
- Local coding helpers
- Cost-sensitive RAG generation
Strengths
- Excellent quality-per-parameter
- Easy local deployment (Ollama / LM Studio)
- Permissive licensing on many releases
Limitations
- Not a substitute for frontier models on hardest tasks
- Multimodal coverage depends on specific Phi variant
Related guides
Related benchmarks
Related research
Related GitHub
Related tools
Related rankings
Companies
Explore more models
- Llama 4MetaMeta’s Llama 4 family — open-weight multimodal models designed for research and commercial use under Meta’s community license.
- DeepSeek V3DeepSeekDeepSeek’s MoE general model — strong open-weight performance on coding and knowledge tasks with competitive API pricing.
- MixtralMistralMistral’s sparse Mixture-of-Experts open models (e.g. Mixtral 8x7B / 8x22B) — efficient high-quality text generation for self-hosting.
- Qwen3AlibabaAlibaba’s Qwen3 generation — strong multilingual open models spanning chat, reasoning modes, coding, and multimodal variants.
- DeepSeek R1DeepSeekDeepSeek’s reasoning-focused model trained with reinforcement learning for multi-step math, science, and coding problem solving.
- Claude OpusAnthropicAnthropic’s highest-capability Claude tier for deep reasoning, long-context analysis, coding, and careful instruction following.