Products & Services · · Team LATENT
Claude Sonnet 5.5 launches with more than 30% faster output
Claude Sonnet 5.5 speeds up output while keeping API rates unchanged. We compare performance and costs in evaluations from Artificial Analysis, Vals AI, Mercor, and others, and explain availability and API migration considerations.
Anthropic has announced Claude Sonnet 5.5. The company says it generates output more than 30% faster than its predecessor and reduces the cost per task. What improved without a change in API rates? We examine the comparisons and practical changes.
Key points
- The developer reports cost reductions of up to 30%. However, some independent evaluations find that maximum reasoning makes it more expensive than Opus 5.5. API rates and total execution costs need separate consideration.
- It surpasses GPT-6 Astra on an independent composite index. Coding, knowledge work, and image understanding also improve over its predecessor, with some published scores above Opus 5.5. Rankings still vary by field and evaluation conditions.
What happened
- Anthropic announced Claude Sonnet 5.5 on September 28. It is the second Claude 5.5 model, following Opus 5.5.
- It is available through Claude apps, the API, AWS, Google Cloud, and Microsoft Azure.
- The API model name is
claude-sonnet-5-5. Rates are unchanged from Sonnet 5.
Primary sources: Anthropic's announcement, pricing
The developer reports up to 30% lower cost per task
Anthropic says Sonnet 5.5 completes the same work with fewer tokens. In its tests, cost per task was up to 30% lower than Sonnet 5. This is the developer's measurement, not a promise that every request will cost 30% less. The official announcement also reports more than 30% faster output generation.
Output speed and time to complete a request are different measures. Reading files and waiting for external services also take time, so 30% faster output cannot be interpreted as finishing work 30% sooner.
A practical cost comparison should include all work through completion, including retries. Cheap intermediate answers may not lower the total if they require more corrections. Alongside speed, the result needs to meet the requested quality.
It surpasses Opus 5.5 on terminal-based development tasks
In coding evaluations included in Anthropic's announcement, Sonnet 5.5 outperformed Sonnet 5. Improvements appear in Terminal-Bench 4.0, which involves multiple steps using a terminal, and CursorBench 4.0, whose tasks derive from actual Cursor usage.
Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0, exceeding Opus 5.5's 66.4% by 4.2 percentage points. According to page 113 of the official evaluation document, Sonnet 5.5 used maximum reasoning, max, while Opus 5.5 used xhigh, its highest-scoring setting. Both results average five trials per task. The published score is higher, but reasoning settings differ and measurements have variability.
Values published by Anthropic. Sonnet 5.5 uses max; Opus 5.5 uses xhigh on Terminal-Bench and max on CursorBench.
Measurement conditionsAnthropic measured Terminal-Bench; Cursor measured CursorBench and supplied the results to Anthropic. Scores come from the announcement's comparison table. Reasoning and execution settings appear on pages 113 and 115 of the official evaluation document. Terminal-Bench includes safety-related fallback to an older model. LATENT did not rerun these evaluations.
Opus 5.5 remains positioned above it for complex, open-ended work
On CursorBench 4.0, Opus 5.5 leads with 57.8%, compared with Sonnet 5.5's 55.5%. Anthropic positions Sonnet 5.5 for clearly scoped everyday tasks and says Opus 5.5 remains stronger on complex work requiring sustained judgment. Leading on Terminal-Bench does not imply an advantage on every task.
One way to compare their roles is to separate small bug fixes from design work that reconsiders requirements. Scores alone cannot determine how much of your own work can be delegated.
It also surpasses GPT-6 Astra on an independent composite index
Sonnet 5.5 scored 56 on the Artificial Analysis composite index. At each model's maximum reasoning setting, the displayed scores were 58 for Opus 5.5, 53 for GPT-6 Astra, and 48 for GPT-6 Sol. Sonnet trails Opus but leads both OpenAI models.
The index combines ten evaluations, including knowledge work, scientific programming, and long-context understanding. According to the evaluator's announcement, Sonnet 5.5 scored 71% on AutomationBench-AA, which measures workflow automation, compared with Opus 5.5's 70%. On factual knowledge benchmark AA-Omniscience, Opus led with 66% accuracy against 54%. A higher composite score does not imply superiority in every capability.
However, this evaluation used a prerelease model with a bug that degraded structured-output responses. Anthropic says it fixed the bug at release, and Artificial Analysis plans to rerun affected evaluations. Fallback to older models was also enabled for Claude; Sonnet 5.5 fell back to Sonnet 5 on about 0.1% of tasks.
Knowledge work approaches Opus 5.5, with a higher APEX score
In Artificial Analysis evaluations of work products, Sonnet 5.5 scored 1844 on GDPval-AA v2.1 and 1811 on AA-Briefcase v1.1. Opus 5.5 scored 1846 and 1822, respectively, leaving small gaps. These are Elo ratings calculated from comparisons of work products, not accuracy percentages.
On Mercor's APEX-Agents, Sonnet 5.5 scored 75.5% against Opus 5.5's 73.5%. The evaluation tests whether models complete 240 investment banking, management consulting, and corporate law tasks using multiple apps. Both used max reasoning, with Sonnet 5.5 leading on the published score.
On APEX-SWE's 200 software development tasks, however, Sonnet 5.5 scored 66.4% against Opus 5.5's 67.6%. Mercor's reported uncertainty intervals overlap, so the small difference should not be treated as decisive.
Strong scores in code migration, app building, and biological reasoning
In Vals AI's evaluations, Sonnet 5.5 scored 69.22% on the composite Vals Index, behind Opus 5.5's 69.69%. Its component scores included 69.83% on Code Migration, 92.39% on app-building benchmark Vibe Code Bench v1.1, and 81.11% on BioMysteryBench. Each led its respective ranking when checked.
Its strengths are not uniform. MedScribe, which covers medical document writing, scored 91.10%, while Harvey's Legal Agent Benchmark scored 2.92%, near the bottom. Different benchmarks use different tasks and scoring methods, so their percentages cannot be compared directly.
Vals AI generally used max reasoning and fell back to Sonnet 5 when the model refused an answer. Counting tasks helped by fallback as failures reduces the Vals Index to 68.89%. Fallback has a particularly large effect on cybersecurity and incident-response evaluations.
Chart reading and computer use improve over its predecessor
Anthropic reports that Chartography, which tests specialist chart reading, rose from Sonnet 5's 15.6% to 61.6% for Sonnet 5.5. Opus 5.5 scored 64.4%. All are measured without tools.
On computer-use benchmark OSWorld 2.1, Sonnet 5.5 scored 80.1%, above Sonnet 5's 57.0% and close to Opus 5.5's 81.8%. This score includes partial credit for incomplete tasks; it is not the percentage of tasks fully completed.
On Humanity's Last Exam, which spans difficult questions across disciplines, Sonnet 5.5 scored 64.5% with tools, up from Sonnet 5's 54.9%. These figures come from Anthropic's evaluation document and should be distinguished from the independent measurements discussed earlier.
Input and output rates are half those of Opus 5.5
The official pricing table lists Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens. That matches Sonnet 5 and is half the rate for Opus 5.5. Cache reads, which reuse stored input, cost $0.20 per million tokens for all three models.
U.S. dollars per million tokens, comparing standard API use.
Input
Sonnet 5$2
Sonnet 5.5$2
Opus 5.5$4
Output
Sonnet 5$10
Sonnet 5.5$10
Opus 5.5$20
Cache reads
Sonnet 5$0.20
Sonnet 5.5$0.20
Opus 5.5$0.20
Official pricingAnthropic's pricing table and announcement, checked September 29. Cache writes, tools, and other features may incur additional charges.
Half the rate does not necessarily mean half the total cost
If input and output token counts are identical, those charges are half those of Opus 5.5. In practice, models produce different response lengths and take different steps. Cache-read rates are also identical, so the price table alone does not justify assuming half the total cost.
In Artificial Analysis's composite-index tasks, Sonnet 5.5 at maximum reasoning cost $7.60 per task, exceeding Opus 5.5's $5.98. GPT-6 Astra cost $3.26 and GPT-6 Sol $1.06. Sonnet 5.5 used many tokens to reach its high score, so maximum reasoning is not always economical.
Meanwhile, on Vals AI's Vals Index, Sonnet 5.5 cost $20.80 per test against Opus 5.5's $32.77. The tasks and execution methods differ from Artificial Analysis, so dollar amounts from the two sites cannot be compared as equivalent measures.
In Box's evaluation of complex enterprise work, Sonnet 5.5 achieved 65% accuracy, up from Sonnet 5's 61%. Box reports roughly 2.4 times faster completion of work products and 12% lower total token use. These results reflect Box's own environment. Anthropic's cost-reduction claims and independent findings of higher costs at maximum reasoning concern different work and conditions.
GitHub Copilot availability has begun
GitHub also announced general availability in Copilot on September 28. It says prerelease testing achieved coding results similar to Sonnet 5 with fewer steps, tokens, and tool calls, and shorter completion times. These are GitHub's reported usage results, different in nature from an independent ranking under common conditions.
Eligible plans include Copilot Pro, Pro+, Max, Business, and Enterprise. The model can be selected in VS Code, Visual Studio, Copilot CLI, GitHub apps, and other interfaces. The rollout is gradual, so it may not appear immediately. Organizational use also depends on administrator model policies.
API migration changes how reasoning is disabled
Google Cloud's model specifications list text, image, and PDF input with text output. The input limit is one million tokens and the output limit is 128,000 tokens. Accepting many documents does not guarantee using every detail, so long-document workflows should also check references and supporting evidence.
As explained in Anthropic's migration guide, switching an existing API integration requires checking reasoning settings as well as changing the model name. Sonnet 5.5 does not support thinking.type: disabled. Workflows that previously disabled reasoning are directed to use between_tools instead. This mode skips extended initial reasoning and returns brief progress text between tool calls as reasoning blocks.
Programs expecting only answer text should check this response format. Alongside limits and speed, migration should verify that conversations and tool execution continue to work in the existing app.
High-risk cyber tasks can fall back to the older model
Anthropic describes a fallback to Sonnet 5 for high-risk cyber requests. The API migration guide says the model stops answering affected requests and retries with Sonnet 5 when server-side fallback is enabled. The API does not always switch automatically. Anthropic says ordinary software development, including finding and fixing bugs, remains available.
Sources and references
- Introducing Claude Sonnet 5.5 Anthropic / September 28, 2026
- Claude API pricing Anthropic / Checked September 29, 2026
- Claude Sonnet 5.5 in GitHub Copilot GitHub / September 28, 2026
- Claude Sonnet 5.5 on Google Cloud Google Cloud / Checked September 29, 2026
- Claude Sonnet 5.5 System Card Anthropic / Evaluation conditions and evaluators
- Sonnet 5.5 migration guide Anthropic / Checked September 29, 2026
- Claude Sonnet 5.5 composite index and execution cost Artificial Analysis / Checked September 29, 2026
- Sonnet 5.5 evaluation and prerelease-model caveats Official Artificial Analysis post on X / Checked September 29, 2026
- Claude Sonnet 5.5 Benchmarks, Cost and Capabilities Vals AI / Checked September 29, 2026
- APEX-Agents and APEX-SWE results Mercor / Checked September 29, 2026
- Sonnet 5.5 APEX results compared with the previous model Official Mercor post on X / Checked September 29, 2026
- Claude Sonnet 5.5 on real enterprise knowledge work Box / Checked September 29, 2026