AI Fundamentals

Claude Models Guide

Engineering guide to Anthropic Claude — Sonnet 5.5, Opus 5.5, Fable 5.1, and Haiku 5.5 tiers, long context, adaptive thinking, MCP, and workload-based production routing.

55 min readIntermediateLast reviewed: 8 October 2026

Quick Summary

Anthropic's Claude family is a production lineup — Sonnet 5.5 for most work, Opus 5.5 for high-stakes escalation, Fable 5.1 as a separate higher-priced SKU, Haiku 5.5 for latency and volume — selected by workload, never as a permanent single default.

One Analogy

Claude tiers are like staffing a project: Haiku is the fast junior for tickets, Sonnet is the senior who ships most features, Opus is the principal for high-stakes design, Fable is the specialist you call for the longest unattended runs — assign by problem, not habit.

Engineering Rule

Never hardcode one Claude tier as the permanent default; route by effort and tier, pin versioned model IDs, and eval-gate upgrades and route-downs.

TL;DR

  • Current GA tiers: Claude Sonnet 5.5 (claude-sonnet-5-5, 2026-09-28) — everyday production workhorse at the same $2/$10 list price as Sonnet 5; Claude Opus 5.5 (claude-opus-5-5, 2026-09-22) — high-stakes reasoning and agentic work at $4/$20 per 1M tokens; Claude Fable 5.1 (claude-fable-5-1) — separate higher-priced SKU ($10/$50) that Anthropic says Opus 5.5 matches on most work; Claude Haiku 5.5 (claude-haiku-5-5, 2026-10-07) — latency, volume, and subagents at $0.10/$0.50 per 1M tokens for prompts up to 100K ($0.50/$2.50 above). Route by workload — never one permanent default.

  • Context: 1M tokens and 128K max output on Sonnet 5.5; ~1M on Opus 5.5 / Fable 5.1 (Fable 5.1 also documents 128k max output). Haiku 5.5 also has 1M context and 128K max output, with higher per-token prices above 100K-token prompts (verify per model in Anthropic docs). Long context is a real product differentiator for codebase and document packs — with cost and attention caveats.

  • Adaptive / effort thinking — raise thinking effort for hard paths; keep low effort (or Haiku) for volume. On Opus 5.5, thinking cannot be disabled. Fable 5.1 defaults to High effort in Claude Code and Medium in Claude apps. As of 2026-09-16, chat and Cowork merge into one conversation (Pro/Max first); Claude Docs, Slides, and Design can be requested from any chat in beta. Claude Code Projects (beta, 2026-09-17) add a coordinator that delegates to parallel cloud threads. Messages API on-demand compaction is a 2026-09-14 beta (compact-2026-09-04). Tier × effort is the routing grid.

  • Historical strengths: careful instruction following, long-context coding, Constitutional AI alignment story, and Model Context Protocol for tool connectivity.

  • Content marking (2026-08): models launched on or after 2026-08-02 embed SynthID-Text watermarks (EU AI Act transparency, applied globally). C2PA credentials on supported files. Does not identify users; detection API is in private preview. See AI Security.

  • History only: Claude 2 → 3 → 3.5/3.7 → 4.x → Sonnet 5 / Opus 5 / Fable 5.1. Do not start new systems on claude-sonnet-5 or older IDs; pin claude-sonnet-5-5.

Quick Decision Guide

If you want to... Read
Route OpenAI GPT tiers GPT Models
Use Anthropic Claude Claude Models
Use Google Gemini Gemini Models
Self-host open weights Llama · Mistral · DeepSeek
Reduce model cost Cost Optimization
Compare on your own tasks Evaluation

Who this guide is for

  • Best for: AI engineers · ML engineers · backend engineers · architects
  • Difficulty: Intermediate
  • Estimated time: 55 min

Learning Path

Large Language Models → Prompt Engineering → Claude Models → Function Calling → Cost Optimization → Evaluation

On this page

Why This Matters

Anthropic's Claude is the primary peer to OpenAI GPT in many production stacks — IDE agents like Cursor, enterprise copilots, and agent frameworks. The decision is not "Claude vs GPT" as a brand war. It is which tier and thinking effort meet your latency, context, safety, and cost constraints on a measured golden set.

Teams that default every request to Opus overpay. Teams that force Haiku onto multi-file refactors underperform. Understanding Claude's tier grid, long-context behavior, and honest failure modes lets you route deliberately and avoid vendor hype from either side.

Operationally, Claude forces the same discipline as GPT: pin IDs, attribute tokens (including thinking), and re-eval on every upgrade. The difference is where quality shows up first — long packs, careful edits, MCP-hosted tools — and where friction shows up first — verbosity, over-refusal, and effort-driven latency. Design the product around those realities instead of assuming a single "best model" checkbox.

Engineering Insight

Most production failures trace back to weak routing and evaluation, not to picking the "wrong" provider. Choose tiers by workload and eval-gate every change.

The Problem Claude Models Solve

Enterprise LLM deployments hit three recurring frictions:

  1. Context limits — policies, contracts, and repos span more tokens than early chat models allowed without aggressive chunking.
  2. Unreliable instruction following — models drift, invent policy, or ignore constraints.
  3. Tool integration sprawl — every datasource needs custom connector code.

Claude addresses these with large context on leading tiers, post-training aimed at helpful/honest/harmless behavior (Constitutional AI lineage), and MCP as a standardized tool/data protocol. For developers, Claude often excels when the job is careful reading of long inputs: code review, contract analysis, research synthesis, multi-file edits.

You still need RAG, tools, structured outputs, and evaluation. Long context is not a substitute for retrieval design or claim verification — see hallucinations and large language models.

How We Got Here

Diagram: Claude family evolution

timeline
    title From Claude 2 chat to Fable 5.1
    2023 : Claude 2
         : Early long-context chat alternative
    2024 : Claude 3 / 3.5
         : Haiku / Sonnet / Opus product grid
    2024-2025 : 3.7 / early 4.x
              : Stronger coding + extended thinking
    2025-2026 : Sonnet 5 / Opus 5 / Haiku 4.5
              : ~1M context class + effort routing
    2026-09 : Fable 5.1, Opus 5.5, Sonnet 5.5
            : Fable Sep 1; Opus Sep 22 at $4/$20; Sonnet Sep 28 at $2/$10
    2026-10 : Haiku 5.5
            : Oct 7; first Haiku with an effort setting

The Haiku / Sonnet / Opus naming still describes everyday production routing. Opus 5.5 is the current Opus id. Fable 5.1 remains a separate higher-priced SKU.

Era Representative models Engineering lesson
Claude 2 Claude 2 Viable GPT alternative; smaller ecosystem
Claude 3.x Haiku 3, Sonnet 3.5, Opus 3 Clear speed/quality/cost grid
3.7 / early 4.x Sonnet 4, Opus 4 Coding + extended thinking
Current Sonnet 5.5, Opus 5.5, Fable 5.1, Haiku 5.5 Workhorse / high-stakes / higher-priced SKU / volume

Keep Claude 2–4.x (including Haiku 4.5), claude-opus-5, and claude-sonnet-5 in migration history. New systems should pin claude-sonnet-5-5 / claude-opus-5-5 / Fable 5.1 / claude-haiku-5-5 from Anthropic docs.

What Is the Claude Model Family?

Claude is Anthropic's family of decoder-only transformer LLMs, trained and aligned with techniques that include Constitutional AI — models critique/revise against written principles during training — plus RLHF-style preference optimization.

Tier Model ID (typical) Role
Haiku claude-haiku-5-5 Fast, inexpensive volume, subagents, and low-latency UX (2026-10-07)
Sonnet claude-sonnet-5-5 Default production workhorse (2026-09-28, same $2/$10 list price as Sonnet 5)
Opus claude-opus-5-5 High-stakes reasoning, coding, and agentic workloads ($4/$20)
Fable claude-fable-5-1 Separate higher-priced SKU ($10/$50); Anthropic says Opus 5.5 matches it on most work

Claude Sonnet 5.5 (claude-sonnet-5-5, 2026-09-28) is the current Sonnet API model. Anthropic prices it at $2 input / $10 output per 1M tokens, the same list price as Sonnet 5, with cache reads at $0.10 (halved from $0.20 on 2026-10-07) and 5-minute cache writes at $2.50. It documents a 1M context window and 128K max output. Adaptive thinking is on by default; between_tools turns off up-front thinking. Forced tool use returns an error. On the Claude API and Google Cloud, computer_20251124 is not accepted. Claude Opus 5.5 (claude-opus-5-5, 2026-09-22) is the current Opus API model at $4/$20, cache reads at $0.20, cache writes at $5, and fast mode at $8/$40. Anthropic says it matches Fable 5.1 on most work and costs less to run than Opus 5. Thinking cannot be disabled, and forced tool use returns an error. Fable 5.1 (claude-fable-5-1) remains available at $10/$50, with cache reads at $0.25. Mythos 5.1 is the trusted-access twin of Fable, not of Opus. Prefer Sonnet 5.5 for most production traffic and Opus 5.5 when Sonnet fails evals. Pin Fable 5.1 only when an eval shows a gap that justifies the higher list price. Claude Haiku 5.5 (claude-haiku-5-5, 2026-10-07) is priced by prompt length: $0.10 input / $0.50 output per 1M tokens up to 100K tokens, and $0.50 / $2.50 above that, with cache reads at $0.01 ($0.05 above 100K). It has a 1M context window, 128K max output, and adaptive thinking with default effort medium, and it is the first Haiku with an effort setting. Its tokenizer counts about 30% more tokens than Haiku 4.5 for the same text, so compare cost per task, not per token. Anthropic positions it for summaries, compaction, classification, and subagents; Sonnet 5.5 and Opus 5.5 remain better for complex agentic coding.

Additional product surfaces (availability varies — verify):

  • Adaptive / extended thinking — internal reasoning tokens before the visible answer; on Sonnet 5.5, between_tools is the setting that keeps up-front thinking off
  • Computer use — GUI interaction via screenshots/actions (where offered)
  • Batch API — discounted async jobs
  • Prompt caching — discounted repeated prefixes
  • MCP — standardized tool and context servers — Model Context Protocol

Note

Model IDs, context limits, and prices change. Verify against Anthropic docs and the current pricing page before architecture or finance lock-in.

How Claude Models Work

Claude tokenizes text (and images on vision-enabled models), runs transformer attention across the context window, and generates tokens autoregressively until stop. The application owns everything around that loop: retrieval packing, tool execution, schema validation, and delivery policy. Treating Claude as a black-box "smart API" without those layers reproduces the same failure modes as any other LLM family.

Long context. Sonnet 5.5 / Opus 5.5 class windows (~1M) let you pass large repos or document sets in one request. That is genuinely useful when cross-file or cross-clause dependencies matter and chunking would sever them. You still pay for every input token and can hit "lost in the middle" — put critical instructions at the edges; measure faithfulness on long packs with questions that probe middle sections. Haiku 5.5 also has a 1M window, but prompts over 100K tokens cost five times more per token, so keep high-volume Haiku routes short. Whole-repo dumps are not free intelligence: if your eval shows RAG + Sonnet matching whole-pack quality at half the cost, prefer RAG.

Thinking / effort. Extended or adaptive thinking allocates internal tokens before the user-visible answer. Higher effort helps math, planning, and hard coding; it increases latency and cost. Route effort + tier together: low-effort Sonnet or Haiku for FAQ; high-effort Sonnet or Opus for hard agents. Log thinking-token counts separately so finance and SRE can see which routes burn hidden compute. Do not ship a global "max effort" flag — that is the Opus-for-everything mistake in another form.

Tool use / MCP. Claude emits tool calls similar to function calling. MCP standardizes how hosts discover and call tools — useful for IDE and multi-server agent setups. Keep authorization outside the model: Claude proposing run_sql does not mean the host should execute unrestricted queries. Sandbox tools, scope credentials, and cap agent loops.

Alignment behavior. Constitutional AI lineage shows up as careful refusals and instruction adherence. That reduces some unsafe completions; it can also over-refuse legitimate edge cases — design UX for "blocked → clarify / escalate." Measure refusal rate by intent class so safety wins do not silently become product regressions.

Prompt caching and Batch. Structure prompts with static policy, style, and corpus prefixes first; put user-specific variables last so cache keys remain stable across turns. Use Batch for offline review queues and nightly synthesis where multi-hour SLA is acceptable — typically at a meaningful discount versus interactive tokens (verify current Anthropic Batch pricing).

Architecture

Diagram: Claude production architecture

flowchart TB
    subgraph App [Application]
        U[User / Agent host]
        R[Router: risk × complexity]
    end
    subgraph Claude [Anthropic API]
        H[claude-haiku-5-5]
        S[claude-sonnet-5-5]
        O[claude-opus-5-5]
    end
    subgraph Controls [Controls]
        Eff[Thinking effort]
        Cache[Prompt cache]
        MCP[MCP / tools]
        Schema[Structured outputs]
        Eval[Eval + traces]
    end
    U --> R
    R -->|volume| H
    R -->|default product| S
    R -->|hardest| O
    S --> Eff
    O --> Eff
    H --> Cache
    S --> MCP
    O --> MCP
    MCP --> Schema --> Eval

Sonnet 5.5 carries most traffic; Opus 5.5 is the usual hard escalation; Fable 5.1 is the higher-priced SKU; Haiku absorbs volume.

Component Responsibility
Router Map intent/risk → Haiku / Sonnet / Opus / Fable + effort
Pinned ID Snapshot per environment; never silent "latest" in prod
MCP / tools Live data and actions outside parametric memory
Prompt cache Stable system + corpus prefix; variables last
Eval Golden coding/doc packs before tier or effort changes

Lineup snapshot (October 2026)

Model Best for Context (approx.) Relative cost Latency
claude-haiku-5-5 Volume, classify, subagents 1M; 128K max output $0.10/$0.50 up to 100K Fastest (default effort medium)
claude-sonnet-5-5 Most production RAG, coding, agents 1M; 128K max output $2/$10; cache read $0.10 Fast (API default effort high)
claude-opus-5-5 High-stakes reasoning / agentic ~1M class — verify $4/$20; cache read $0.20 Variable
claude-fable-5-1 Separate higher-priced SKU ~1M class; 128k max out $10/$50; cache read $0.25 Variable (High effort default in Claude Code)

Pricing changes frequently — verify Anthropic's published rates (and cache/batch discounts) rather than locking finance to blog numbers. Treat cost like GPT: route, cache, Batch, compress — cost optimization.

Step-by-Step Flow

Diagram: Claude request with effort routing

sequenceDiagram
    participant U as User
    participant App as App / IDE host
    participant Rt as Router
    participant C as Claude
    participant T as Tools / MCP
    U->>App: Task
    App->>Rt: Score difficulty + stakes
    Rt-->>App: tier + thinking effort
    App->>C: Messages + tools + effort
    alt tool / MCP call
        C-->>App: tool_use
        App->>T: Invoke
        T-->>App: Result
        App->>C: tool_result
    end
    C-->>App: Final text / structured
    App-->>U: Deliver + log tokens

Decide tier and effort before the first token; escalate on validation failure, not by habit.

  1. Define context shape — whole-repo pack vs RAG chunks vs short chat.
  2. Choose starting tier — Sonnet 5.5 (claude-sonnet-5-5) for most product paths; Haiku for volume; Opus 5.5 (claude-opus-5-5) when Sonnet fails evals; Fable 5.1 only when an eval shows a gap that justifies $10/$50.
  3. Set thinking effort — low for FAQ; raise on Sonnet/Opus for hard planning.
  4. Pin model IDs in config; record in traces.
  5. Wire tools/MCP for live systems; schemas for parsers — structured outputs.
  6. Enable prompt caching on large static prefixes.
  7. Instrument tokens (including thinking), latency, cost, refusals.
  8. Eval-gate tier and effort changes on coding and faithfulness suites — evaluation.
Workload Start Escalate
FAQ / classify Haiku 5.5 Sonnet if confidence low
RAG product chat Sonnet 5.5 Opus on multi-hop fail
Multi-file coding Sonnet 5.5 + higher effort Opus 5.5
Long-horizon agents Opus 5.5 Fable 5.1 if an eval shows a gap
Long doc synthesis Sonnet 5.5 (1M) Opus for hardest synthesis
High-volume extract Haiku + Batch Sonnet if quality dips

Real Production Example

Code-review assistant: Haiku triages diff size/risk; Sonnet reviews; Opus only for flagged architectural diffs.

from __future__ import annotations

import os
from anthropic import Anthropic

client = Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])

HAIKU = os.environ.get("CLAUDE_VOLUME_MODEL", "claude-haiku-5-5")
SONNET = os.environ.get("CLAUDE_MODEL", "claude-sonnet-5-5")
OPUS = os.environ.get("CLAUDE_COMPLEX_MODEL", "claude-opus-5-5")


def triage_diff(diff_text: str) -> str:
    msg = client.messages.create(
        model=HAIKU,
        max_tokens=256,
        messages=[
            {
                "role": "user",
                "content": (
                    "Classify this git diff as low, medium, or high risk. "
                    "Reply with one word only.\n\n" + diff_text[:12000]
                ),
            }
        ],
    )
    return msg.content[0].text.strip().lower()


def review_diff(diff_text: str, risk: str) -> str:
    model = OPUS if risk.startswith("high") else SONNET
    # Raise thinking/effort via current API params when available — verify docs.
    msg = client.messages.create(
        model=model,
        max_tokens=4096,
        messages=[
            {
                "role": "user",
                "content": (
                    "Review this diff for correctness, security, and API breaks. "
                    "Cite file hunks. If unsure, say so.\n\n" + diff_text
                ),
            }
        ],
    )
    return msg.content[0].text


diff = open("change.diff").read()
risk = triage_diff(diff)
print("risk:", risk)
print(review_diff(diff, risk))

For agent hosts, prefer MCP servers for repo, browser, and DB tools instead of one-off HTTP wrappers — see MCP and function calling.

Design Decisions

Choose Claude when:

  • Long-context coding or document analysis is central and Sonnet/Opus win your bake-off
  • You want careful instruction following and a strong refusal posture
  • MCP-centric tooling (IDE agents, multi-server hosts) fits your architecture
  • Prompt caching + Batch make large static corpora economical

Prefer another family when:

  • OpenAI ecosystem / Azure networking is already standard and OpenAI's current flagship GPT models win evals → GPT
  • GCP grounding, Vertex IAM, or native multimodal video is the center → Gemini
  • You must self-host → Llama / Mistral

Whole-pack context vs RAG is the recurring Claude-specific decision. Use ~1M packs when the task needs global connectivity (cross-file refactors, multi-contract consistency). Use RAG when the corpus is mostly independent chunks and you can prove recall@k. Many teams hybridize: retrieve candidates, then stuff a ranked subset into Sonnet — cheaper than naive whole-corpus prompts, richer than tiny top-3 packs.

Common patterns

Pattern Practice
Sonnet-first Default product model; Opus on escalation only
Effort dial Low effort volume; high effort hard agents
Whole-pack vs RAG Use ~1M when graph connectivity matters; else retrieve
Cache-heavy corpus Legal/codebase prefix cached across turns
MCP host One protocol for filesystem, DB, browser tools
Shadow Opus Sample Sonnet failures offline on Opus to tune escalation rules

Diagram: Tier × effort decision

flowchart TD
    Q[Task] --> Vol{High volume / low stakes?}
    Vol -->|Yes| H[Haiku 5.5]
    Vol -->|No| Hard{Fails Sonnet eval or max difficulty?}
    Hard -->|No| S[Sonnet 5.5]
    Hard -->|Yes| O[Opus 5.5]
    S --> E{Need deeper thinking?}
    E -->|Yes| SH[Sonnet + higher effort]
    E -->|No| SL[Sonnet + low effort]

Escalate tier and effort independently; measure both on your golden set.

For IDE and agent hosts, prefer MCP servers with least-privilege credentials over embedding long-lived cloud keys in the model prompt. The model should request tools; the host should authorize them. That boundary matters more for Claude-heavy coding agents than for simple chat wrappers.

Comparisons

Dimension Claude GPT Gemini Llama
Current workhorses Sonnet 5.5 / Opus 5.5 / Fable 5.1 / Haiku 5.5 GPT-6 Astra / Sol / Luna Google Flash / Pro-class Size-dependent
Context (approx.) ~1M across current tiers — verify Large — verify Large Flash — verify Often smaller
Tool calling Tools + MCP Mature tools/schemas Tools + Search grounding DIY
Volume tier Haiku 5.5 OpenAI volume tier Flash Self-host
Alignment story Constitutional AI lineage RLHF / policy stack Google policy stack You align
Typical enterprise plane Anthropic API / cloud partners Azure OpenAI common Vertex AI Your GPUs
Claude tier Prefer for Avoid as permanent default for
Haiku 5.5 Latency, classify, subagents Hard multi-file agents
Sonnet 5.5 Most production traffic Pure spam at massive QPS (try Haiku)
Opus 5.5 High-stakes reasoning/agents Everyday FAQ
Fable 5.1 Peak long-horizon agents Default product chat

Bake-offs should fix the task pack first (coding edits, long-doc QA, tool agents), then vary only the model ID and effort. Changing prompts and models at once makes winners meaningless. Report quality, p95 latency, and $ per successful task — not isolated arena scores.

Common Mistakes

  1. Opus for everything — burns budget; Sonnet 5.5 should carry most load.
  2. Max thinking effort globally — latency and cost explode; dial per route.
  3. Dumping megatokens without measurement — long context ≠ perfect recall; eval middle-span questions.
  4. Ignoring over-refusal — build clarify/escalate paths for blocked legitimate asks.
  5. Unpinned aliases — pin IDs; re-eval on every upgrade.
  6. Skipping cache layout — dynamic timestamps in the system prompt destroy prompt-cache hits.
  7. No peer bake-off — Claude is not automatically best; compare GPT and Gemini on your suite.

Where It Breaks Down

  • Knowledge cutoff — RAG/tools for current facts.
  • Verbosity — Claude can over-explain; constrain with prompts and max tokens.
  • Over-refusal — safety posture blocks some medical/legal/security research UX.
  • Latency under high effort — stream UI; move hard jobs async.
  • Multimodal gaps vs Gemini — for heavy video/audio pipelines, evaluate Gemini honestly.
  • Vendor lock-in via MCP-only assumptions — keep tool interfaces portable where possible.
  • Cache invalidation surprises — editing a "static" policy prefix busts prompt-cache hit rates overnight; version the cached blob explicitly.
  • Thinking-token opacity — product UIs that hide thinking still bill for it; finance must see those meters.

When NOT to Default to Claude

Do not set Claude (or Opus) as the permanent org default when:

  • OpenAI or Gemini win your latency/cost/quality Pareto on the real workload
  • You need GCP-native grounding and Vertex controls as the system of record
  • Almost all traffic is cheap classification — Haiku or another volume model after bake-off
  • You cannot yet log thinking tokens and $ — instrument first
  • Residency requires self-host open weights

Warning

Permanent Opus (or permanent max effort) is an anti-pattern. Route by capability, latency, and cost; pin versions; re-eval on change.

Running in Production

Best Practice

Default to Sonnet 5.5, escalate to Opus 5.5 (claude-opus-5-5) on measured failure, use Fable 5.1 only when an eval shows a gap, absorb volume and subagent work on Haiku 5.5, and treat thinking effort as a per-route dial. On Opus 5.5, thinking cannot be disabled. On Sonnet 5.5, between_tools turns off up-front thinking.

Dimension Guidance
Scaling Stateless API; watch TPM/RPM; queue Haiku for bulk
Cost Tier routing + prompt cache + Batch — cost optimization
Latency Haiku / low effort for chat UX; Opus async for hard jobs
Security Server-side keys; tool sandboxing; audit MCP servers; expect SynthID-Text on new Claude outputs (AI security)
Observability Tier, effort, thinking tokens, TTFT, refusals, tool errors
Evaluation Coding + faithfulness golden sets; CI on ID changes — evaluation
Reliability Backoff on 429; fallback to GPT/Gemini twin route

Long-context packs deserve their own SLO: track p95 latency and cost per review job separately from interactive chat. A Sonnet 5.5 whole-repo review that is acceptable asynchronously can destroy interactive TTFT budgets if you reuse the same route. Split configs: CLAUDE_INTERACTIVE_MODEL vs CLAUDE_BATCH_MODEL, with different max_tokens and effort defaults.

Continue Learning

Production Checklist

  • Model ID pinned per environment (Haiku / Sonnet / Opus / Fable tiers)
  • Thinking effort dial set per route (not a global max)
  • Interactive vs long-pack / Batch configs separated
  • Prompt caching enabled on large static prefixes
  • MCP/tools scoped with least-privilege credentials
  • SynthID-Text / C2PA marking understood for compliance and detector workflows
  • Structured outputs enforced on parser paths
  • Token and cost metrics include thinking tokens
  • Golden-set gates for coding and faithfulness upgrades
  • Fallback provider documented and tested
  • Refusal / escalation UX defined
  • Rollback strategy for tier or ID changes documented

Prerequisites

Core Concepts

Implementation

Optimization

Advanced Topics

Diagram: Claude learning path

flowchart LR
    LLM[LLMs] --> CL[Claude models]
    CL --> CW[Context windows]
    CW --> MCP[MCP]
    MCP --> FC[Function calling]
    FC --> CO[Cost opt]
    CO --> EV[Evaluation]
    CL --> GPT[GPT]
    CL --> GM[Gemini]

Long context and MCP sit next to Claude; compare peers before locking a vendor.

Interview Questions

  1. How do Sonnet 5.5, Opus 5.5, Fable 5.1, and Haiku 5.5 differ?
    Sonnet 5.5 (claude-sonnet-5-5, $2/$10) is the production workhorse; Opus 5.5 (claude-opus-5-5) is the current high-stakes escalation at $4/$20; Fable 5.1 is a separate $10/$50 SKU that Anthropic says Opus 5.5 matches on most work; Haiku 5.5 (claude-haiku-5-5, from $0.10/$0.50) is volume, latency, and subagent work. Route by workload.

  2. When do you raise thinking effort?
    On hard planning, math, or multi-step agents — after measuring that low effort fails. Not on every FAQ.

  3. Is 1M context a reason to skip RAG?
    Sometimes for tightly coupled packs; usually no. Cost, latency, and middle-context loss still favor retrieval for large corpora.

  4. What is MCP?
    A protocol for connecting models to tools and data sources with standard discovery/invocation — see MCP.

  5. Claude vs GPT for coding?
    Bake off Sonnet 5.5 / Opus 5.5 / Fable 5.1 vs GPT-6 Astra and Sol on your repo metrics — do not assume a permanent winner.

  6. How does Constitutional AI show up in products?
    Stronger refusal and principle-following tendencies; plan for over-refusal UX.

  7. How do you control Claude cost?
    Haiku for volume, cache prefixes, Batch offline, escalate Opus rarely — cost optimization.

  8. Why pin model IDs?
    Reproducible evals and traces; aliases can change behavior without a deploy.

Key Takeaways

  • Claude's production grid is Sonnet 5.5 / Opus 5.5 / Fable 5.1 / Haiku 5.5 — route by tier and thinking effort. On Opus 5.5, thinking cannot be disabled. On Sonnet 5.5, between_tools turns off up-front thinking. Pin claude-sonnet-5-5 for new systems.
  • Long context is powerful and expensive; measure middle-span faithfulness.
  • MCP and tools ground actions; schemas and eval ground reliability.
  • Historical Claude versions belong in migration notes, not new defaults.
  • Never treat Opus, Fable, or Claude generally as the permanent org-wide default.

FAQs

Is Sonnet 5.5 always the right default?

It is the right starting default for many product apps, not a permanent law. If Haiku meets evals, use it; if Sonnet fails hard tasks, escalate to Opus. Pin claude-sonnet-5-5.

When should I use Opus 5.5?

When Sonnet 5.5 fails your golden set on hard reasoning, architecture, or long agent loops. Pin claude-opus-5-5 (2026-09-22, $4/$20 per 1M tokens). Thinking cannot be disabled, and forced tool use returns an error. Escalate to Fable 5.1 only when an eval shows a gap that justifies $10/$50.

When should I use Fable 5.1?

When an eval shows Opus 5.5 is not enough for the job, and the higher list price is justified. Anthropic says Opus 5.5 matches Fable 5.1 on most work. Pin claude-fable-5-1. Mythos 5.1 is the trusted-access twin of Fable, not a public default. Do not assume Fable is cheaper than Opus 5.5; Opus 5.5’s cache reads are $0.20 per 1M tokens and Fable 5.1’s are $0.25.

What is Haiku 5.5 for?

Latency-sensitive and high-volume paths: classification, extraction, summaries, compaction, triage, cheap first-pass routing, and subagents under a Sonnet or Opus lead. Pin claude-haiku-5-5 (2026-10-07). Anthropic lists Haiku 4.5 for retirement not sooner than 2026-10-15, so migrate and re-run evals; the newer tokenizer changes token counts.

How large is Claude's context?

1M context and 128K max output on Sonnet 5.5 and Haiku 5.5; ~1M class on Opus 5.5 / Fable 5.1. Haiku 5.5 charges more per token for prompts over 100K. Verify current model pages.

Does Claude eliminate hallucinations?

No. Alignment and long context help; you still need grounding and verification.

Claude vs Gemini for documents?

Claude is strong on careful long-text reasoning; Gemini is strong on native multimodal and Google grounding. Bake off on your corpus.

Should I use prompt caching?

Yes for large static system/doc prefixes. Keep dynamic fields at the end. For long agent threads, Anthropic also offers on-demand Messages API compaction (beta header compact-2026-09-04, 2026-09-14): you choose when to summarize prior turns and replay a signed compaction block. That is not the same as OpenAI Agents API automatic session compaction.

Can I still pick Chat vs Cowork?

On the new Claude apps experience (rolling out to Pro and Max first, 2026-09-16), you do not pick a mode — Claude routes tools from one conversation. Existing Cowork tasks, projects, connectors, and files carry over. Until your account has the new UI, Chat/Cowork selectors still work as before.

What are Claude Code Projects?

As of 2026-09-17, Projects in Claude Code are a coordinator conversation: Claude scopes the work, delegates to parallel threads (each a Claude Code cloud session on its own branch), reviews outputs, and shares memory plus a project library. Beta for select Pro/Max cloud-session users with no existing web/desktop projects; broader Claude Code, then chat/Cowork and Team/Enterprise later. Threads are cloud-only at launch. Distinct from Cursor Projects.

Can I fine-tune Claude?

Anthropic's fine-tune options vary by program and time — check current docs; most teams start with prompting, tools, and RAG.

Are Claude outputs watermarked?

Models launched on or after 2026-08-02 embed SynthID-Text watermarks (applied globally). The mark is not user-identifying. Supported files can carry C2PA credentials. Anthropic’s detection API is in private preview for eligible organizations. See How Claude’s text watermark works.

References

Further Reading

Next Topics

Learning Path

Continue Learning

Related Guides

Related companies

  • Anthropic

    Enterprise-first AI company focused on safe, reliable reasoning models.

Related models

  • Claude Opus

    Anthropic’s Claude Opus 5.5 (API id claude-opus-5-5, released 2026-09-22) for complex agentic coding, enterprise work, long-context analysis, and careful instruction following. Anthropic says it matches Fable 5.1 on most work at a lower price. Fable 5.1 remains a separate higher list-price SKU. Sonnet 5.5 (claude-sonnet-5-5, 2026-09-28) is the current Sonnet at the same $2/$10 list price as Sonnet 5. Haiku 5.5 (claude-haiku-5-5, 2026-10-07) is the current Haiku.

  • Claude Sonnet

    Anthropic’s Claude Sonnet 5.5 (API id claude-sonnet-5-5, released 2026-09-28) for well-scoped everyday coding, agents, and documents. List price matches Sonnet 5 at $2/$10 per 1M tokens. Opus 5.5 remains the high-stakes tier; Haiku 5.5 (2026-10-07) is the low-cost subagent and high-volume tier.

  • Claude Fable

    Anthropic’s Claude Fable 5.1 — the most capable widely released Claude for long-horizon agents, deep reasoning, and demanding coding workflows. Mythos 5.1 is the same underlying model with relaxed cyber/life-sciences safeguards for trusted-access programs.

  • Claude Haiku

    Anthropic’s fast, cost-efficient Claude tier for high-volume chat, classification, extraction, and sub-agent steps where latency and price matter more than peak reasoning. Claude Haiku 5.5 (API id claude-haiku-5-5, released 2026-10-07) is the current Haiku: Anthropic’s fastest model at standard speed, the first Haiku with an adjustable effort setting, and around 75% cheaper to run than Haiku 4.5 on average. Sonnet 5.5 and Opus 5.5 remain better for complex agentic coding.

Related Tools

ToolCategoryPurposeWebsiteBest For
Claude
Featured
ai productsAnthropic’s conversational AI focused on reliability and safety.claude.aiLong document analysis
Claude Code
TrendingAPICloud
codingTerminal-first coding agent from Anthropic with long-context codebase reasoning, Projects coordinators (beta), /design artboards, and CLI/IDE/desktop surfaces.claude.aiTerminal-first agentic coding
Cursor
TrendingAPICloud
codingAI-native code editor with codebase context, multi-file agents, Projects (persistent coordinators), Origin code hosting, self-hosted Cloud Agent machines, and intelligent model routing for teams.cursor.comAI-native IDE development