rkj dev

Anthropic Releases Claude Haiku 5.5 With Effort Controls and Lower Pricing

The lightweight model arrives with major agentic performance gains, adjustable effort controls, and immediate deployment across AWS, Google Cloud, and Azure.

Modern enterprise data center with illuminated optical server racks
AI-generated illustration

Key takeaways

  • Anthropic states the model costs around 75% less on average to run compared to Haiku 4.5, pricing standard prompts up to 100,000 tokens at $0.10 per million input tokens.
  • The release marks the first Haiku model with adjustable effort settings, enabling developers to balance speed, cost, and reasoning depth.
  • Anthropic launched the model simultaneously on the Claude Platform, Amazon Bedrock, Google Cloud, and Microsoft Azure on October 7, 2026.
  • Companion updates include a 50% price cut for Claude Sonnet 5.5 cache reads and new monthly API credits for Claude Max and Team subscribers.

On October 7, 2026, Anthropic announced Claude Haiku 5.5, introducing the latest edition of its lightweight, fast model family. Billed as the fastest and most efficient model in the Claude 5.5 lineup, Haiku 5.5 delivers substantial benchmark improvements across agentic coding, computer use, and knowledge extraction. Anthropic reported that the model costs roughly 75% less to run on average than Claude Haiku 4.5, directly targeting high-volume workflows, repetitive enterprise operations, and real-time subagent routing.

The launch immediately rolled out across major cloud platforms. Alongside direct API access, Claude Haiku 5.5 is generally available on Amazon Web Services via Amazon Bedrock, Google Cloud, and Microsoft Azure, with AWS Bedrock and Google Cloud backing the deployment with retirement commitments extending no earlier than October 7, 2027.

Software developer reviewing performance metrics and latency telemetry on an office display
AI-generated illustration

Technical Specs and the Addition of Effort Controls

Claude Haiku 5.5 includes a 1-million-token context window and supports up to 128,000 maximum output tokens, according to technical documentation published by AWS Documentation and Google Cloud. Supported input modalities include text, images, and documents such as PDFs, with text output capabilities.

For the first time in the Haiku product tier, Anthropic introduced configurable effort controls. Adaptive thinking is active by default and can be calibrated across low, medium, high, xhigh, and max levels, defaulting to medium. If reasoning is disabled, the effort setting is capped at high. This functionality allows developers to adjust the balance between reasoning depth and execution costs per task rather than applying a blanket standard across an entire system.

In testing documented by Simon Willison, a low-effort prompt generating SVG graphics completed in 7 seconds at a cost of 0.0936 cents, while a max-effort run with expanded reasoning traces took 5 minutes and 9 seconds, costing 3.3826 cents. Willison also observed that Haiku 5.5 adopts a modified tokenizer, which processes approximately 1.25 times as many tokens for identical prompts compared to Haiku 4.5.

Benchmark Performance Across Agentic Tasks

Anthropic highlighted significant performance gains over Claude Haiku 4.5 and competitive positioning against OpenAI's GPT-6 Luna across evaluation suites outlined in the Anthropic announcement:

  • Computer Use (OSWorld 2.1 offline subset): Haiku 5.5 achieved 72.4%, up from 15.7% on Haiku 4.5, surpassing GPT-6 Luna's 48.9% and approaching Sonnet 5.5's 83.9%.
  • Agentic Coding (Terminal-Bench 4.0): Haiku 5.5 scored 39.2%, compared to 0.0% for Haiku 4.5 and 16.4% for GPT-6 Luna. Claude Sonnet 5.5 posted 70.6% on the same benchmark.
  • Multidisciplinary Reasoning (Humanity's Last Exam): Haiku 5.5 recorded 45.9% without tools and 57.4% with tools, improving over Haiku 4.5's 10.2% without tools and 18.7% with tools.
  • Knowledge Work (GDPval-AA v2.1): The model reached 1,620, higher than Haiku 4.5 (735) and GPT-6 Luna (1,437).
  • Visual Reasoning (Chartography without tools): Haiku 5.5 reached 46.4%, up from 6.4% on Haiku 4.5 and 29.1% on GPT-6 Luna.

Anthropic shared testing feedback from enterprise software firm Asana. Staff software engineer Aaron Vinh reported that integrating Haiku 5.5 into Asana's AI Teammates product line resulted in an inference speedup of up to 2.5 times per agent turn and reduced task-completion latency by over 30% during bug triaging and project portfolio monitoring.

Abstract diagram of primary reasoning hub delegating work to smaller subagent streams
AI-generated illustration

Tiered Pricing and Ecosystem Updates

Anthropic structured the pricing for Claude Haiku 5.5 around prompt length thresholds, reported by Anthropic and analyzed by Simon Willison:

  • Prompts up to 100,000 tokens: $0.10 per million input tokens, $0.50 per million output tokens, $0.01 per million cache reads, and $0.125 per million cache writes.
  • Prompts exceeding 100,000 tokens: $0.50 per million input tokens, $2.50 per million output tokens, $0.05 per million cache reads, and $0.625 per million cache writes.

Anthropic stated that roughly 90% of requests handled by the previous Haiku model fell under the 100,000-token threshold. For comparison, Haiku 4.5 was priced flatly at $1.00 per million input tokens, $5.00 per million output tokens, $0.10 per million cache reads, and $1.25 per million cache writes.

Concurrent with the release, Anthropic halved the cache read pricing for Claude Sonnet 5.5 from $0.20 down to $0.10 per million tokens. The company stated this change lowers the total operational expense of running Sonnet 5.5 on most multi-turn agentic workflows by around 20%.

Additionally, Anthropic introduced recurring monthly API credits for subscribers using the Claude Platform. Max 5x subscribers receive $100 per month, Max 20x subscribers receive $200 per month, and Team tier organizations receive up to $500 per month pooled across members. These credits correspond to the monthly subscription fees, do not roll over month-to-month, and can be applied toward any model on the platform.

Anthropic also released beta updates to its official Python and TypeScript software development kits, adding direct integration for computer use and browser automation workflows.

Safety Safeguards and Cloud Availability

Evaluations outlined by Anthropic show fewer instances of misaligned behaviors and reduced willingness to cooperate with misuse compared to Haiku 4.5. On cybersecurity boundaries, Anthropic stated that Haiku 5.5 enforces stricter rules than Haiku 4.5 while permitting more defensive diagnostic operations than Sonnet 5.5. Active penetration testing tools and techniques favored by attackers remain blocked. Safeguards surrounding biological research mirror Sonnet 5, Sonnet 5.5, and Opus 5, with expanded access reserved for verified entities enrolled in Anthropic's Cyber and Life Sciences verification programs.

According to AWS and 9to5Mac, Haiku 5.5 is designed to function effectively alongside Claude Opus 5.5 and Sonnet 5.5. Under this architectural pattern, larger models manage overarching task decomposition, system planning, and code review, while fleets of parallel Haiku 5.5 subagents execute focused classification, file transformations, database queries, and document compaction.

On Amazon Bedrock, developers can deploy Haiku 5.5 via regional, geographic cross-region, and global inference profiles, with full support on AWS GovCloud endpoints. On Google Cloud, the model is available generally through Agent Studio and Model Garden under multi-region and global endpoints with processing limits up to 30 million input tokens per minute.

Frequently asked questions

How does Claude Haiku 5.5 pricing compare to Haiku 4.5?

For prompts up to 100,000 tokens, Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, representing a 90% price reduction from Haiku 4.5 ($1.00 input and $5.00 output per million tokens), while Anthropic states the model costs around 75% less to run on average overall across workloads.

What are effort controls on Claude Haiku 5.5?

Effort controls allow developers to configure reasoning depth across low, medium, high, xhigh, and max levels. Medium is the default setting, enabling teams to balance response latency and cost against task complexity.

Where is Claude Haiku 5.5 currently available?

Claude Haiku 5.5 is available on the Claude Platform API, Amazon Web Services via Amazon Bedrock, Google Cloud via Agent Studio and Model Garden, and Microsoft Azure.

What adjustments were made to Claude Sonnet 5.5 pricing?

Anthropic cut the cost of prompt cache reads for Claude Sonnet 5.5 by 50%, reducing the rate from $0.20 to $0.10 per million tokens, which decreases overall agentic workload costs by approximately 20%.

Sources

  1. Introducing Claude Haiku 5.5anthropic.com · Official
  2. Introducing Claude Haiku 5.5 on AWS | Amazon Web ServicesAmazon Web Services · Oct 7, 2026 · Official
  3. www-cdn.anthropic.comwww-cdn.anthropic.com · Official
  4. Claude Haiku 5.5 - Amazon Bedrockdocs.aws.amazon.com · Official
  5. Claude Haiku 5.5 on Google CloudGoogle Cloud Documentation · Official
  6. Claude Haiku 5.5Simon Willison’s Weblog
  7. Anthropic upgrades Claude with new Haiku 5.5 model, details here9to5Mac · Oct 7, 2026

How this story was made: the newsroom picked it up from Hacker News, Google News and Techmeme, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (29 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.

#Anthropic #Claude Haiku 5.5 #Amazon Bedrock #Google Cloud #AI Benchmarks #Cloud Computing

Published October 8, 2026 at 00:59 UTC