AI news, with the context that matters.

Products & Services · ·

Claude Opus 5.5 launches with roughly 60% lower cost per evaluation task

Claude Opus 5.5 improves its score in an independent evaluation while reducing cost per task by about 60%. We explain performance, pricing and practical caveats.

Anthropic has released Claude Opus 5.5. Input and output rates fall 20%, yet an independent evaluation finds roughly 60% lower cost per task. What contributes beyond the price cut? We examine the results and cost breakdown to identify when savings occur.

Key takeaways

  1. Lower unit prices and reduced token usage produce a higher evaluation score at lower cost than the predecessor.
  2. Maximum reasoning raises the score further, but costs about the same as the predecessor at that setting.
  3. Input and output rates are 40% of Fable 5.1's. At API defaults, its evaluation score is slightly below Fable 5.1 at maximum reasoning.
Official Claude Opus 5.5 announcement image
Source: Anthropic

What happened

  • Anthropic announced Claude Opus 5.5 at around 1:30 a.m. Japan time on September 23. API, AWS, Google Cloud and Microsoft Azure availability began that day.
  • Rates are $4 per million input tokens and $20 per million output tokens, down from $5 and $25 for Opus 5.
  • It is the first release in the 5.5 series. Sonnet 5.5 and Haiku 5.5 are expected within weeks.
  • About 90 minutes later, OpenAI announced GPT-6 Sol and Luna. Sol costs $2 for input and $10 for output; Luna costs $0.10 and $0.50.

Primary sources: Anthropic announcement, Pricing, Artificial Analysis evaluation

A higher evaluation score at lower cost per task

Independent evaluator Artificial Analysis published results on announcement day. Its composite indexNote 1 rises from 48 for Opus 5 to 51 for Opus 5.5. Average cost per evaluation task falls from $3.61 to $1.34, a reduction of about 60%.

The price cut is not the only cause. Opus 5.5 produces about 26,000 output tokensNote 2 per task, a little over half the predecessor's roughly 46,000, while achieving a higher score.

Opus lets users specify how much computation to devote to reasoning before answering, referred to here as reasoning effortNote 3. More effort can help with difficult questions but also takes time and money. Both results above use API defaults. Opus 5.5 defaults to one level lower effort than its predecessor, yet scores higher.

Opus 5 needed maximum effort to score the same 51, at $5.86 per task—more than four times Opus 5.5's cost. This is an aggregate comparison; individual tests vary. Opus 5.5 leads on terminal-based development, but trails Opus 5 at maximum effort on two knowledge-work tests.

The same evaluation allows comparison with other models. At maximum effort, Fable 5.1 scores 53 at $7.63 per task; OpenAI's GPT-6 Astra also scores 53 at $3.26. Default Opus 5.5 is slightly below both, but increasing effort one level raises its score to 54 at $1.82. The index methodology changed in early September, so these values cannot be compared with those in our GPT-6 Astra launch article.

Figure 1Evaluation scores and cost per task by model

Bar lengths show US dollars per evaluation task. The index combines multiple tests.

API defaultsOpus 5.5 uses medium effort; Opus 5 uses high

Claude Opus 5.5New model

Index51

Cost$1.34

Claude Opus 5Predecessor

Index48

Cost$3.61

At maximum effort

Claude Opus 5.5New model

Index58

Cost$5.98

Claude Opus 5Predecessor

Index51

Cost$5.86

Claude Fable 5.1High-end model

Index53

Cost$7.63

GPT-6 AstraOpenAI

Index53

Cost$3.26

At API defaults, Opus 5.5 scores higher and costs about 60% less. At maximum effort, its cost is almost the same as its predecessor's.

Independent measurementsMeasured by Artificial Analysis. Safety fallback to other models is enabled for Opus 5.5 and Fable 5.1. Indexes are rounded to whole numbers and costs to two decimals. Evaluation article, Model page (September 22, 2026).

At the highest score, cost is almost unchanged from the predecessor

The score increases with effort, reaching 58 at max. This exceeds Fable 5.1 and Astra, but costs $5.98 per task: almost unchanged from Opus 5's $5.86 at max, and about 1.8 times Astra's. At matched effort levels up to one below max, Opus 5.5 is cheaper.

Cost barely changes because token usage rises. At max effort, Opus 5.5 produces roughly 119,000 tokens per task including internal reasoning, about 1.6 times the predecessor's 73,000. Input token usage is also about twice as high. Increased usage offsets lower prices. Every, which published an early review on announcement day, likewise reports substantial token consumption when work is delegated without limits.

Evaluation conditions also matter. Opus 5.5 can switch processing to another model when it detects requests related to cyberattacks and other restricted areas. Artificial Analysis enabled that feature, so some tasks may include answers from a different model.

Input and output prices fall 20%; cache reads fall 60%

Input and output rates fall 20%: from $5 to $4 and $25 to $20 per million tokens. CacheNote 4 reads drop 60%, from $0.50 to $0.20. Anthropic says cache reads account for most costs in agentic work and coding.

Anthropic says typical tasks cost 40% less than Opus 5 without an explicit effort setting, but does not disclose the workload mix. The editors therefore calculated a separate example: ten million input tokens, 90% read from cache, plus 200,000 output tokens cost $9.80 instead of $14.50, about 32% less. This excludes cache writes. At the same token counts without caching, the reduction is 20%.

Achieving a 40% reduction under those assumptions requires reduced token usage as well as lower prices. In the independent evaluation above, Opus 5.5 achieves a higher score using fewer tokens than its predecessor.

Anthropic's range increases in capability and price from Haiku to Sonnet to Opus, with Fable and MythosNote 5 above them. Anthropic says Opus 5.5 matches Fable 5.1 on most work. Fable 5.1 costs $10 for input and $50 for output, making Opus 5.5's rates 40% as high.

Figure 2Pricing comparison

US dollars per million tokens. Models differ in capability. GPT-6 Sol is included for pricing only, not a performance comparison.

Anthropic

Claude Fable 5.1High-end model

Input$10

Output$50

Claude Opus 5Predecessor

Input$5

Output$25

Claude Opus 5.5New model

Input$4

Output$20

Claude Sonnet 5Lower-tier model

Input$2

Output$10

OpenAI

GPT-6 AstraAnnounced in September

Input$10

Output$50

GPT-6 SolAnnounced the same day

Input$2

Output$10

SourcesAnthropic pricing, OpenAI model page, TechCrunch (GPT-6 Sol).

In Anthropic's announcement, Opus 5.5 beats Fable 5.1 on every directly comparable test. Yet the company characterizes it overall as matching Fable on most work, saying score differences at this level do not translate cleanly into practical gaps and internal experience shows less difference. One Every reviewer estimates its coding ability at about 90% of Fable's; others in the same article describe it as equal or occasionally better.

Customer reports describe shorter task times

Anthropic also presents examples from users and companies testing long-running work. These are favorable reports selected by the developer and may not represent general performance.

One tester completed an audit and repair of 200,000 lines of code in under three hours, compared with more than 20 hours and 2.5 times the tokens for Opus 5. Another completed a 680,000-line migration in under a day. In Anthropic's own experiment, acquisition analysis fell from 93 to 63 minutes at half the cost.

There are also reports of unattended work. Clio says it ran for over 18 hours without intervention. Quantium reports that a complex coding task previously requiring 38 prompts over four days took three hours and 11 prompts. Optiver reports roughly half the steps, time and output to reach Opus 5 quality on agentic coding tasks, with costs 40–50% lower.

Figure 3Changes in task time and prompt counts

Use cases published in Anthropic's announcement.

  • Audit and repair 200,000 lines of codeEarly-user report

    Opus 5Over 20 hours

    Opus 5.5Under 3 hours

  • Complex coding workReported by customer Quantium

    Previously4 days, 38 prompts

    Opus 5.53 hours, 11 prompts

  • Acquisition analysisAnthropic experiment

    Opus 593 min

    Opus 5.563 min

  • Code migration from C to RustAnthropic internal test; cost down 51%

    Fable 5.112 hours

    Opus 5.59.5 hours

  • Unattended workReported by customer Clio

    Opus 5.5Over 18 hours without intervention

User reportsExamples from early users, Quantium and Clio, selected by Anthropic for its announcement.

Developer measurementsAcquisition analysis and code migration. Anthropic announcement (September 22, 2026).

Fewer prompts can reduce the burden of intervening, correcting or requesting a restart. Whether similar gains carry over to writing documents or conducting research needs to be tested on actual work.

Some requests are answered by older models

Some items in the announcement trail other models. On business automation, GPT-6 Astra scores 41.4% versus Opus 5.5's 40.0%; on scientific-research tasks, Astra scores 64.6% versus 58.7%. According to Anthropic's notes, the automation results were measured and published by Zapier: Astra's comes from its public leaderboard and Opus 5.5's from early-access evaluation. Astra's science score is a figure published by OpenAI.

The responding model can change with the request. Many cyber tasks, including exploit generation and penetration testing, route to the older Opus 4.8. Some biological tasks and a narrow subset of frontier-AI development, such as low-level software for certain AI chips, route to Opus 5. Anthropic says most ordinary AI development is unaffected. Source-code vulnerability discovery and writing secure code can remain on Opus 5.5. Apps switch automatically; in the API, affected requests stop unless the user enables fallback.

Anthropic reports safety evaluations across roughly 2,000 scenarios and says nearly every measure is its best among recent Claude models. However, it also often observed the model suspecting it was being tested. Behavior could differ between evaluation and deployment, so these results alone cannot establish safety.

External evaluator METR concludes that AI research-and-development capability is slightly above Fable 5.1's but full automation remains difficult. The report explicitly says Anthropic had an opportunity to review and edit it.

Switching through the API requires checking existing code

One approach is to try actual work at API defaults, then increase effort if needed. Even if the previous model used high effort, compare results before carrying that setting forward. In the evaluation above, max effort costs more than four times the default.

In API applications, changing only the model name may not be enough. Four breaking changes affect existing code, including errors for disabling reasoning or forcing certain tool use. Anthropic's migration guide lists the changes.

Access from Japan through the app, API and clouds

Users in Japan can access it through Claude's app and API, AWS, Google Cloud and Microsoft clouds. App access covers paid Pro, Max, Team and Enterprise plans; the free plan is not listed. Five-hour limits also increase for Pro, Max, Team and seat-based Enterprise, though the announcement does not quantify the increase.

For Japan-only processing, Amazon Bedrock offers a Japan endpoint, with processing confined to Tokyo and Osaka regions. Anthropic's pricing lists a 10% premium for regional processing. Zero-data-retention arrangementsNote 6 are supported. Higher-tier Fable 5.1 normally retains data for 30 days, with zero-retention access limited to eligible enterprises.

The announcement does not describe Japanese-language performance. Users should test their actual Japanese documents and conversations.

Timeline

Dates are in US time.

  1. Anthropic releases Mythos Preview to a limited audience, its first demonstration of a model above Opus.

  2. Fable 5 and Mythos 5 are announced: Fable for broad availability and Mythos for restricted access.

  3. Opus 5 is announced at $5 input and $25 output.

  4. Fable 5.1 and Mythos 5.1 are announced.

  5. OpenAI announces GPT-6 Astra. Its president argues that pricing should be assessed per task. Our GPT-6 Astra launch article

  6. Anthropic's CEO publishes an essay calling for a slower pace of development.

  7. Opus 5.5 is announced, Anthropic's first release since the call for restraint. It is September 23 in Japan.

Sources and references

Update history

  • Published.
  • Added the distinction between Zapier's public leaderboard and early-access evaluation for AutomationBench, following the notes in Anthropic's announcement.
  • Corrected the blanket statement that ordinary AI development is excluded from model fallback to reflect the developer's wording.

LATENT uses Google Analytics (Firebase) to improve our articles. With your permission, cookies and similar technologies send browsing and interaction data to Google. Privacy policy