Foundation Models
Canonical reference pages for major LLMs and multimodal models. Each model links to benchmarks, guides, tools, research, and rankings.
21 models · Research feed · Benchmarks · LLMs guide
Featured Models
GPT-5.6
OpenAIOpenAI’s GPT-5.6 family (Sol flagship, Terra balanced, Luna cost-efficient) for complex reasoning, coding, multimodal understanding, and agentic tool use. The gpt-5.6 API alias routes to Sol.
GPT-6 Astra
OpenAIOpenAI’s GPT-6 Astra peak model (API id gpt-6-astra) for computer use, coding agents, professional artifacts, science, and cybersecurity-adjacent defender workflows. Staged launch 2026-09-03; GPT-5.6 Sol/Terra/Luna remain the volume and cost-routing family.
Claude Opus
AnthropicAnthropic’s Claude Opus 5 tier for complex agentic coding, enterprise work, long-context analysis, and careful instruction following. Claude Fable 5.1 sits above Opus for peak widely released capability.
Claude Sonnet
AnthropicAnthropic’s Claude Sonnet 5 tier — best combination of speed and intelligence for most production agents and coding, at lower cost than Opus.
Claude Fable
AnthropicAnthropic’s Claude Fable 5.1 — the most capable widely released Claude for long-horizon agents, deep reasoning, and demanding coding workflows. Mythos 5.1 is the same underlying model with relaxed cyber/life-sciences safeguards for trusted-access programs.
Claude Haiku
AnthropicAnthropic’s fast, cost-efficient Claude tier for high-volume chat, classification, extraction, and sub-agent steps where latency and price matter more than peak reasoning.
Gemini 3.1 Pro
GoogleGoogle’s current Pro-class Gemini for hard reasoning and native multimodal work. Prefer API id gemini-3.1-pro-preview; Gemini 3.5 Pro remains partner-testing. Legacy gemini-2.5-pro is scheduled for shutdown Oct 16, 2026.
DeepSeek V3
DeepSeekDeepSeek’s MoE general model — strong open-weight performance on coding and knowledge tasks with competitive API pricing.
DeepSeek R1
DeepSeekDeepSeek’s reasoning-focused model trained with reinforcement learning for multi-step math, science, and coding problem solving.
DeepSeek V4
DeepSeekDeepSeek’s V4 generation — deepseek-v4-pro (V4-Pro-0813 GA) and deepseek-v4-flash (Flash-0731) with 1M context, thinking effort low/high/max, native Responses API, and strong agentic coding. Experimental multimodal API: deepseek-v4-flash-vision-exp (2026-08-21).
Kimi K3
Moonshot AIMoonshot’s Kimi K3 — 2.8T MoE (104B active) open-weight multimodal agentic model with 1M context, native vision, and strong long-horizon coding. Weights on Hugging Face under the Kimi K3 License.
Muse Spark
MetaMeta Superintelligence Labs’ Muse Spark 1.3 — closed multimodal reasoning model for agentic tasks, long-horizon coding, computer use, and 1M-context workflows via Muse Code and the Meta Model API. Succeeds Spark 1.2; max reasoning is still withheld for safety testing.
Muse Glimmer
MetaMeta Superintelligence Labs’ Muse Glimmer — Apache-2.0 ~30B dense multimodal agent model for on-device and single-GPU local agents. Sibling to closed Muse Spark; distinct from Llama 4.
Qwen3
AlibabaAlibaba’s Qwen3 family spanning Qwen3.8-Max (2.4T MoE / 95B active; live cloud alias qwen3.8-max routes to the 0902 snapshot as of 2026-09-05), open Qwen3.8-27B (dense VLM, Apache-2.0), and Qwen3.8-Flash-Next (125B / 6B active multimodal MoE + 51B n-gram embeddings)—a Qwen4 architecture preview for cost-efficient agentic coding. Production Qwen3.8-Flash on QwenCloud adds 1M-default context and built-in tools atop the Flash-Next design.
Llama 4
MetaMeta’s Llama 4 family — open-weight multimodal models designed for research and commercial use under Meta’s community license.
Proprietary
GPT-5.6
OpenAIOpenAI’s GPT-5.6 family (Sol flagship, Terra balanced, Luna cost-efficient) for complex reasoning, coding, multimodal understanding, and agentic tool use. The gpt-5.6 API alias routes to Sol.
GPT-6 Astra
OpenAIOpenAI’s GPT-6 Astra peak model (API id gpt-6-astra) for computer use, coding agents, professional artifacts, science, and cybersecurity-adjacent defender workflows. Staged launch 2026-09-03; GPT-5.6 Sol/Terra/Luna remain the volume and cost-routing family.
Claude Opus
AnthropicAnthropic’s Claude Opus 5 tier for complex agentic coding, enterprise work, long-context analysis, and careful instruction following. Claude Fable 5.1 sits above Opus for peak widely released capability.
Claude Sonnet
AnthropicAnthropic’s Claude Sonnet 5 tier — best combination of speed and intelligence for most production agents and coding, at lower cost than Opus.
Claude Fable
AnthropicAnthropic’s Claude Fable 5.1 — the most capable widely released Claude for long-horizon agents, deep reasoning, and demanding coding workflows. Mythos 5.1 is the same underlying model with relaxed cyber/life-sciences safeguards for trusted-access programs.
Claude Haiku
AnthropicAnthropic’s fast, cost-efficient Claude tier for high-volume chat, classification, extraction, and sub-agent steps where latency and price matter more than peak reasoning.
Gemini 3.1 Pro
GoogleGoogle’s current Pro-class Gemini for hard reasoning and native multimodal work. Prefer API id gemini-3.1-pro-preview; Gemini 3.5 Pro remains partner-testing. Legacy gemini-2.5-pro is scheduled for shutdown Oct 16, 2026.
Gemini Flash
GoogleGoogle’s Gemini 3.8 Flash workhorse — fast, token-efficient multimodal model for agentic workflows, coding, and high-throughput apps where latency and cost matter. Succeeds 3.7 Flash (GA Aug 2026); 3.7 remains supported for efficiency-first workloads.
Muse Spark
MetaMeta Superintelligence Labs’ Muse Spark 1.3 — closed multimodal reasoning model for agentic tasks, long-horizon coding, computer use, and 1M-context workflows via Muse Code and the Meta Model API. Succeeds Spark 1.2; max reasoning is still withheld for safety testing.
Mistral Large
MistralMistral’s flagship large model for enterprise reasoning, multilingual chat, and function calling via La Plateforme and cloud partners.
Grok
xAIxAI’s Grok 4.6 — frontier coding and long-running agentic model (API id grok-4.6) with 500K context, vision, and strong tool use. Available via the xAI API, Grok Build, Cursor, and partners such as OpenRouter.
Command R+
CohereCohere’s Command R+ model optimized for retrieval-augmented generation, enterprise search, and multilingual business assistants.
Open Source
DeepSeek V3
DeepSeekDeepSeek’s MoE general model — strong open-weight performance on coding and knowledge tasks with competitive API pricing.
DeepSeek R1
DeepSeekDeepSeek’s reasoning-focused model trained with reinforcement learning for multi-step math, science, and coding problem solving.
DeepSeek V4
DeepSeekDeepSeek’s V4 generation — deepseek-v4-pro (V4-Pro-0813 GA) and deepseek-v4-flash (Flash-0731) with 1M context, thinking effort low/high/max, native Responses API, and strong agentic coding. Experimental multimodal API: deepseek-v4-flash-vision-exp (2026-08-21).
Kimi K3
Moonshot AIMoonshot’s Kimi K3 — 2.8T MoE (104B active) open-weight multimodal agentic model with 1M context, native vision, and strong long-horizon coding. Weights on Hugging Face under the Kimi K3 License.
Muse Glimmer
MetaMeta Superintelligence Labs’ Muse Glimmer — Apache-2.0 ~30B dense multimodal agent model for on-device and single-GPU local agents. Sibling to closed Muse Spark; distinct from Llama 4.
Qwen3
AlibabaAlibaba’s Qwen3 family spanning Qwen3.8-Max (2.4T MoE / 95B active; live cloud alias qwen3.8-max routes to the 0902 snapshot as of 2026-09-05), open Qwen3.8-27B (dense VLM, Apache-2.0), and Qwen3.8-Flash-Next (125B / 6B active multimodal MoE + 51B n-gram embeddings)—a Qwen4 architecture preview for cost-efficient agentic coding. Production Qwen3.8-Flash on QwenCloud adds 1M-default context and built-in tools atop the Flash-Next design.
Llama 4
MetaMeta’s Llama 4 family — open-weight multimodal models designed for research and commercial use under Meta’s community license.
Mixtral
MistralMistral’s sparse Mixture-of-Experts open models (e.g. Mixtral 8x7B / 8x22B) — efficient high-quality text generation for self-hosting.
Phi
MicrosoftMicrosoft’s Phi family of small language models — high capability per parameter for on-device, edge, and cost-sensitive deployments.
Reasoning
GPT-5.6
OpenAIOpenAI’s GPT-5.6 family (Sol flagship, Terra balanced, Luna cost-efficient) for complex reasoning, coding, multimodal understanding, and agentic tool use. The gpt-5.6 API alias routes to Sol.
GPT-6 Astra
OpenAIOpenAI’s GPT-6 Astra peak model (API id gpt-6-astra) for computer use, coding agents, professional artifacts, science, and cybersecurity-adjacent defender workflows. Staged launch 2026-09-03; GPT-5.6 Sol/Terra/Luna remain the volume and cost-routing family.
Claude Opus
AnthropicAnthropic’s Claude Opus 5 tier for complex agentic coding, enterprise work, long-context analysis, and careful instruction following. Claude Fable 5.1 sits above Opus for peak widely released capability.
Claude Sonnet
AnthropicAnthropic’s Claude Sonnet 5 tier — best combination of speed and intelligence for most production agents and coding, at lower cost than Opus.
Claude Fable
AnthropicAnthropic’s Claude Fable 5.1 — the most capable widely released Claude for long-horizon agents, deep reasoning, and demanding coding workflows. Mythos 5.1 is the same underlying model with relaxed cyber/life-sciences safeguards for trusted-access programs.
Gemini 3.1 Pro
GoogleGoogle’s current Pro-class Gemini for hard reasoning and native multimodal work. Prefer API id gemini-3.1-pro-preview; Gemini 3.5 Pro remains partner-testing. Legacy gemini-2.5-pro is scheduled for shutdown Oct 16, 2026.
DeepSeek V3
DeepSeekDeepSeek’s MoE general model — strong open-weight performance on coding and knowledge tasks with competitive API pricing.
DeepSeek R1
DeepSeekDeepSeek’s reasoning-focused model trained with reinforcement learning for multi-step math, science, and coding problem solving.
DeepSeek V4
DeepSeekDeepSeek’s V4 generation — deepseek-v4-pro (V4-Pro-0813 GA) and deepseek-v4-flash (Flash-0731) with 1M context, thinking effort low/high/max, native Responses API, and strong agentic coding. Experimental multimodal API: deepseek-v4-flash-vision-exp (2026-08-21).
Kimi K3
Moonshot AIMoonshot’s Kimi K3 — 2.8T MoE (104B active) open-weight multimodal agentic model with 1M context, native vision, and strong long-horizon coding. Weights on Hugging Face under the Kimi K3 License.
Muse Spark
MetaMeta Superintelligence Labs’ Muse Spark 1.3 — closed multimodal reasoning model for agentic tasks, long-horizon coding, computer use, and 1M-context workflows via Muse Code and the Meta Model API. Succeeds Spark 1.2; max reasoning is still withheld for safety testing.
Muse Glimmer
MetaMeta Superintelligence Labs’ Muse Glimmer — Apache-2.0 ~30B dense multimodal agent model for on-device and single-GPU local agents. Sibling to closed Muse Spark; distinct from Llama 4.
Qwen3
AlibabaAlibaba’s Qwen3 family spanning Qwen3.8-Max (2.4T MoE / 95B active; live cloud alias qwen3.8-max routes to the 0902 snapshot as of 2026-09-05), open Qwen3.8-27B (dense VLM, Apache-2.0), and Qwen3.8-Flash-Next (125B / 6B active multimodal MoE + 51B n-gram embeddings)—a Qwen4 architecture preview for cost-efficient agentic coding. Production Qwen3.8-Flash on QwenCloud adds 1M-default context and built-in tools atop the Flash-Next design.
Llama 4
MetaMeta’s Llama 4 family — open-weight multimodal models designed for research and commercial use under Meta’s community license.
Mistral Large
MistralMistral’s flagship large model for enterprise reasoning, multilingual chat, and function calling via La Plateforme and cloud partners.
Grok
xAIxAI’s Grok 4.6 — frontier coding and long-running agentic model (API id grok-4.6) with 500K context, vision, and strong tool use. Available via the xAI API, Grok Build, Cursor, and partners such as OpenRouter.
Command R+
CohereCohere’s Command R+ model optimized for retrieval-augmented generation, enterprise search, and multilingual business assistants.
Phi
MicrosoftMicrosoft’s Phi family of small language models — high capability per parameter for on-device, edge, and cost-sensitive deployments.
Coding
GPT-5.6
OpenAIOpenAI’s GPT-5.6 family (Sol flagship, Terra balanced, Luna cost-efficient) for complex reasoning, coding, multimodal understanding, and agentic tool use. The gpt-5.6 API alias routes to Sol.
GPT-6 Astra
OpenAIOpenAI’s GPT-6 Astra peak model (API id gpt-6-astra) for computer use, coding agents, professional artifacts, science, and cybersecurity-adjacent defender workflows. Staged launch 2026-09-03; GPT-5.6 Sol/Terra/Luna remain the volume and cost-routing family.
Claude Opus
AnthropicAnthropic’s Claude Opus 5 tier for complex agentic coding, enterprise work, long-context analysis, and careful instruction following. Claude Fable 5.1 sits above Opus for peak widely released capability.
Claude Sonnet
AnthropicAnthropic’s Claude Sonnet 5 tier — best combination of speed and intelligence for most production agents and coding, at lower cost than Opus.
Claude Fable
AnthropicAnthropic’s Claude Fable 5.1 — the most capable widely released Claude for long-horizon agents, deep reasoning, and demanding coding workflows. Mythos 5.1 is the same underlying model with relaxed cyber/life-sciences safeguards for trusted-access programs.
Claude Haiku
AnthropicAnthropic’s fast, cost-efficient Claude tier for high-volume chat, classification, extraction, and sub-agent steps where latency and price matter more than peak reasoning.
Gemini 3.1 Pro
GoogleGoogle’s current Pro-class Gemini for hard reasoning and native multimodal work. Prefer API id gemini-3.1-pro-preview; Gemini 3.5 Pro remains partner-testing. Legacy gemini-2.5-pro is scheduled for shutdown Oct 16, 2026.
Gemini Flash
GoogleGoogle’s Gemini 3.8 Flash workhorse — fast, token-efficient multimodal model for agentic workflows, coding, and high-throughput apps where latency and cost matter. Succeeds 3.7 Flash (GA Aug 2026); 3.7 remains supported for efficiency-first workloads.
DeepSeek V3
DeepSeekDeepSeek’s MoE general model — strong open-weight performance on coding and knowledge tasks with competitive API pricing.
DeepSeek R1
DeepSeekDeepSeek’s reasoning-focused model trained with reinforcement learning for multi-step math, science, and coding problem solving.
DeepSeek V4
DeepSeekDeepSeek’s V4 generation — deepseek-v4-pro (V4-Pro-0813 GA) and deepseek-v4-flash (Flash-0731) with 1M context, thinking effort low/high/max, native Responses API, and strong agentic coding. Experimental multimodal API: deepseek-v4-flash-vision-exp (2026-08-21).
Kimi K3
Moonshot AIMoonshot’s Kimi K3 — 2.8T MoE (104B active) open-weight multimodal agentic model with 1M context, native vision, and strong long-horizon coding. Weights on Hugging Face under the Kimi K3 License.
Muse Spark
MetaMeta Superintelligence Labs’ Muse Spark 1.3 — closed multimodal reasoning model for agentic tasks, long-horizon coding, computer use, and 1M-context workflows via Muse Code and the Meta Model API. Succeeds Spark 1.2; max reasoning is still withheld for safety testing.
Muse Glimmer
MetaMeta Superintelligence Labs’ Muse Glimmer — Apache-2.0 ~30B dense multimodal agent model for on-device and single-GPU local agents. Sibling to closed Muse Spark; distinct from Llama 4.
Qwen3
AlibabaAlibaba’s Qwen3 family spanning Qwen3.8-Max (2.4T MoE / 95B active; live cloud alias qwen3.8-max routes to the 0902 snapshot as of 2026-09-05), open Qwen3.8-27B (dense VLM, Apache-2.0), and Qwen3.8-Flash-Next (125B / 6B active multimodal MoE + 51B n-gram embeddings)—a Qwen4 architecture preview for cost-efficient agentic coding. Production Qwen3.8-Flash on QwenCloud adds 1M-default context and built-in tools atop the Flash-Next design.
Llama 4
MetaMeta’s Llama 4 family — open-weight multimodal models designed for research and commercial use under Meta’s community license.
Mistral Large
MistralMistral’s flagship large model for enterprise reasoning, multilingual chat, and function calling via La Plateforme and cloud partners.
Mixtral
MistralMistral’s sparse Mixture-of-Experts open models (e.g. Mixtral 8x7B / 8x22B) — efficient high-quality text generation for self-hosting.
Grok
xAIxAI’s Grok 4.6 — frontier coding and long-running agentic model (API id grok-4.6) with 500K context, vision, and strong tool use. Available via the xAI API, Grok Build, Cursor, and partners such as OpenRouter.
Phi
MicrosoftMicrosoft’s Phi family of small language models — high capability per parameter for on-device, edge, and cost-sensitive deployments.
Vision
GPT-5.6
OpenAIOpenAI’s GPT-5.6 family (Sol flagship, Terra balanced, Luna cost-efficient) for complex reasoning, coding, multimodal understanding, and agentic tool use. The gpt-5.6 API alias routes to Sol.
GPT-6 Astra
OpenAIOpenAI’s GPT-6 Astra peak model (API id gpt-6-astra) for computer use, coding agents, professional artifacts, science, and cybersecurity-adjacent defender workflows. Staged launch 2026-09-03; GPT-5.6 Sol/Terra/Luna remain the volume and cost-routing family.
Claude Opus
AnthropicAnthropic’s Claude Opus 5 tier for complex agentic coding, enterprise work, long-context analysis, and careful instruction following. Claude Fable 5.1 sits above Opus for peak widely released capability.
Claude Sonnet
AnthropicAnthropic’s Claude Sonnet 5 tier — best combination of speed and intelligence for most production agents and coding, at lower cost than Opus.
Claude Fable
AnthropicAnthropic’s Claude Fable 5.1 — the most capable widely released Claude for long-horizon agents, deep reasoning, and demanding coding workflows. Mythos 5.1 is the same underlying model with relaxed cyber/life-sciences safeguards for trusted-access programs.
Claude Haiku
AnthropicAnthropic’s fast, cost-efficient Claude tier for high-volume chat, classification, extraction, and sub-agent steps where latency and price matter more than peak reasoning.
Gemini 3.1 Pro
GoogleGoogle’s current Pro-class Gemini for hard reasoning and native multimodal work. Prefer API id gemini-3.1-pro-preview; Gemini 3.5 Pro remains partner-testing. Legacy gemini-2.5-pro is scheduled for shutdown Oct 16, 2026.
Gemini Flash
GoogleGoogle’s Gemini 3.8 Flash workhorse — fast, token-efficient multimodal model for agentic workflows, coding, and high-throughput apps where latency and cost matter. Succeeds 3.7 Flash (GA Aug 2026); 3.7 remains supported for efficiency-first workloads.
DeepSeek V4
DeepSeekDeepSeek’s V4 generation — deepseek-v4-pro (V4-Pro-0813 GA) and deepseek-v4-flash (Flash-0731) with 1M context, thinking effort low/high/max, native Responses API, and strong agentic coding. Experimental multimodal API: deepseek-v4-flash-vision-exp (2026-08-21).
Kimi K3
Moonshot AIMoonshot’s Kimi K3 — 2.8T MoE (104B active) open-weight multimodal agentic model with 1M context, native vision, and strong long-horizon coding. Weights on Hugging Face under the Kimi K3 License.
Muse Spark
MetaMeta Superintelligence Labs’ Muse Spark 1.3 — closed multimodal reasoning model for agentic tasks, long-horizon coding, computer use, and 1M-context workflows via Muse Code and the Meta Model API. Succeeds Spark 1.2; max reasoning is still withheld for safety testing.
Muse Glimmer
MetaMeta Superintelligence Labs’ Muse Glimmer — Apache-2.0 ~30B dense multimodal agent model for on-device and single-GPU local agents. Sibling to closed Muse Spark; distinct from Llama 4.
Qwen3
AlibabaAlibaba’s Qwen3 family spanning Qwen3.8-Max (2.4T MoE / 95B active; live cloud alias qwen3.8-max routes to the 0902 snapshot as of 2026-09-05), open Qwen3.8-27B (dense VLM, Apache-2.0), and Qwen3.8-Flash-Next (125B / 6B active multimodal MoE + 51B n-gram embeddings)—a Qwen4 architecture preview for cost-efficient agentic coding. Production Qwen3.8-Flash on QwenCloud adds 1M-default context and built-in tools atop the Flash-Next design.
Llama 4
MetaMeta’s Llama 4 family — open-weight multimodal models designed for research and commercial use under Meta’s community license.
Audio
Multimodal
GPT-5.6
OpenAIOpenAI’s GPT-5.6 family (Sol flagship, Terra balanced, Luna cost-efficient) for complex reasoning, coding, multimodal understanding, and agentic tool use. The gpt-5.6 API alias routes to Sol.
GPT-6 Astra
OpenAIOpenAI’s GPT-6 Astra peak model (API id gpt-6-astra) for computer use, coding agents, professional artifacts, science, and cybersecurity-adjacent defender workflows. Staged launch 2026-09-03; GPT-5.6 Sol/Terra/Luna remain the volume and cost-routing family.
Claude Opus
AnthropicAnthropic’s Claude Opus 5 tier for complex agentic coding, enterprise work, long-context analysis, and careful instruction following. Claude Fable 5.1 sits above Opus for peak widely released capability.
Claude Sonnet
AnthropicAnthropic’s Claude Sonnet 5 tier — best combination of speed and intelligence for most production agents and coding, at lower cost than Opus.
Claude Fable
AnthropicAnthropic’s Claude Fable 5.1 — the most capable widely released Claude for long-horizon agents, deep reasoning, and demanding coding workflows. Mythos 5.1 is the same underlying model with relaxed cyber/life-sciences safeguards for trusted-access programs.
Claude Haiku
AnthropicAnthropic’s fast, cost-efficient Claude tier for high-volume chat, classification, extraction, and sub-agent steps where latency and price matter more than peak reasoning.
Gemini 3.1 Pro
GoogleGoogle’s current Pro-class Gemini for hard reasoning and native multimodal work. Prefer API id gemini-3.1-pro-preview; Gemini 3.5 Pro remains partner-testing. Legacy gemini-2.5-pro is scheduled for shutdown Oct 16, 2026.
Gemini Flash
GoogleGoogle’s Gemini 3.8 Flash workhorse — fast, token-efficient multimodal model for agentic workflows, coding, and high-throughput apps where latency and cost matter. Succeeds 3.7 Flash (GA Aug 2026); 3.7 remains supported for efficiency-first workloads.
DeepSeek V4
DeepSeekDeepSeek’s V4 generation — deepseek-v4-pro (V4-Pro-0813 GA) and deepseek-v4-flash (Flash-0731) with 1M context, thinking effort low/high/max, native Responses API, and strong agentic coding. Experimental multimodal API: deepseek-v4-flash-vision-exp (2026-08-21).
Kimi K3
Moonshot AIMoonshot’s Kimi K3 — 2.8T MoE (104B active) open-weight multimodal agentic model with 1M context, native vision, and strong long-horizon coding. Weights on Hugging Face under the Kimi K3 License.
Muse Spark
MetaMeta Superintelligence Labs’ Muse Spark 1.3 — closed multimodal reasoning model for agentic tasks, long-horizon coding, computer use, and 1M-context workflows via Muse Code and the Meta Model API. Succeeds Spark 1.2; max reasoning is still withheld for safety testing.
Muse Glimmer
MetaMeta Superintelligence Labs’ Muse Glimmer — Apache-2.0 ~30B dense multimodal agent model for on-device and single-GPU local agents. Sibling to closed Muse Spark; distinct from Llama 4.
Qwen3
AlibabaAlibaba’s Qwen3 family spanning Qwen3.8-Max (2.4T MoE / 95B active; live cloud alias qwen3.8-max routes to the 0902 snapshot as of 2026-09-05), open Qwen3.8-27B (dense VLM, Apache-2.0), and Qwen3.8-Flash-Next (125B / 6B active multimodal MoE + 51B n-gram embeddings)—a Qwen4 architecture preview for cost-efficient agentic coding. Production Qwen3.8-Flash on QwenCloud adds 1M-default context and built-in tools atop the Flash-Next design.
Llama 4
MetaMeta’s Llama 4 family — open-weight multimodal models designed for research and commercial use under Meta’s community license.
Small Models
Gemini Flash
GoogleGoogle’s Gemini 3.8 Flash workhorse — fast, token-efficient multimodal model for agentic workflows, coding, and high-throughput apps where latency and cost matter. Succeeds 3.7 Flash (GA Aug 2026); 3.7 remains supported for efficiency-first workloads.
Muse Glimmer
MetaMeta Superintelligence Labs’ Muse Glimmer — Apache-2.0 ~30B dense multimodal agent model for on-device and single-GPU local agents. Sibling to closed Muse Spark; distinct from Llama 4.
Mixtral
MistralMistral’s sparse Mixture-of-Experts open models (e.g. Mixtral 8x7B / 8x22B) — efficient high-quality text generation for self-hosting.
Phi
MicrosoftMicrosoft’s Phi family of small language models — high capability per parameter for on-device, edge, and cost-sensitive deployments.