Claude Haiku 5.5
Claude Haiku 5.5 is Anthropic's current small Claude model, released on October 7, 2026. It is designed for high-volume, cost-sensitive work such as classification, summaries, compaction, database queries, browser use, and narrowly scoped coding subagents. The API model ID is claude-haiku-5-5.
What it offers
- Low-cost throughput: for prompts up to 100K tokens, input is $0.10 and output is $0.50 per million tokens. Prompts above 100K cost $0.50 input and $2.50 output per million.
- Prompt-cache pricing: cache reads are $0.01 / $0.05 per million tokens and cache writes are $0.125 / $0.625, for prompts up to / above 100K tokens.
- Adjustable effort: Anthropic says this is the first Haiku-class model with an adjustable effort setting, so teams can trade cost for intelligence.
- Broad availability: Anthropic lists the model as available through the Claude Platform, AWS, Google Cloud, and Microsoft Azure.
When to use it
Use Haiku 5.5 when a workflow has many short, independent calls: triaging support messages, extracting structured fields, routing requests, creating concise summaries, or running a subagent alongside a larger coding model. A practical pattern is to let Claude Sonnet 5.5 plan and review a task while Haiku handles retrieval, compression, or repetitive tool calls.
For an API experiment, start with a small representative batch. Record prompt length, cache-hit rate, output quality, and latency. Then enable the lowest effort level that meets your acceptance test before sending production traffic. This avoids treating vendor benchmark results as a substitute for evaluation on your own data.
Pricing and limits
Anthropic's launch table separates requests at 100K prompt tokens. The lower price is for prompts at or below that threshold; the higher price applies above it. Prices can change, so verify the official announcement and the current API documentation before committing a budget.
| Per 1M tokens | Up to 100K prompt | Over 100K prompt |
|---|---|---|
| Input | $0.10 | $0.50 |
| Output | $0.50 | $2.50 |
| Cache read | $0.01 | $0.05 |
| Cache write | $0.125 | $0.625 |
Limits and alternatives
Haiku 5.5 is not Anthropic's recommended choice for the hardest, long-running agentic coding tasks. The launch announcement continues to position Sonnet 5.5 and Claude Opus 5.5 for more complex work. Its safety controls also still block penetration testing and other higher-risk cyber techniques. Evaluate the model's behavior, tool permissions, and spend controls before giving it access to production systems.
- Claude Sonnet 5.5: better fit when a task needs more sustained agentic coding capability.
- Claude Opus 5.5: use for the most demanding Claude reasoning and coding work.
- GPT-6 Luna: a current low-cost OpenAI option to evaluate alongside Haiku.
FAQ
Is Haiku 4.5 still the current fast Claude model?
No. Haiku 5.5 replaced it as Anthropic's newly released small-model offering. Keep Claude Haiku 4.5 only for a historical comparison.
Where should I start?
Read Anthropic's release notes, run a measured pilot with claude-haiku-5-5, and promote it only after testing output quality and cost on your workload.
Comments
No comments yet. Be the first to comment!
Related Tools
Related Insights

Anthropic Subagent: The Multi-Agent Architecture Revolution
Deep dive into Anthropic multi-agent architecture design. Learn how Subagents break through context window limitations, achieve 90% performance improvements, and real-world applications in Claude Code.
Complete Guide to Claude Skills - 10 Essential Skills Explained
Deep dive into Claude Skills extension mechanism, detailed introduction to ten core skills and Obsidian integration to help you build an efficient AI workflow
The Twilight of Low-Code Platforms: Why Claude Agent SDK Will Make Dify History
A deep dive from first principles of large language models on why Claude Agent SDK will replace Dify. Exploring why describing processes in natural language is more aligned with human primitive behavior patterns, and why this is the inevitable choice in the AI era.