AI news, with the context that matters.

Products & Services · ·

GPT-6.1 Sol launches into a cost contest with Opus 5.5

At half Opus 5.5's standard API rates, GPT-6.1 Sol is a contender for coding and computer use. But the price gap narrows on long inputs, and benchmarks use different conditions. We examine both companies' documentation to assess the trade-offs.

Ask an AI to fix code, read documents, and operate a screen when needed. For repeated work like this, accuracy on a single question is only part of the picture. Can it finish the job? How long will you wait? What does it cost, including retries?

OpenAI's GPT-6.1 Sol, released September 29, is an update aimed squarely at this kind of work. OpenAI describes it as offering capabilities close to the higher-tier GPT-6 Astra at lower cost. Claude Opus 5.5, announced September 22, likewise promises greater efficiency than its predecessor on extended development tasks and knowledge work. OpenAI changelog; GPT-6.1 Sol announcement; Opus 5.5 announcement

Is GPT-6.1 Sol a rival to Opus 5.5? If you are choosing an AI to carry out work and cost matters, it is a serious candidate to compare. Its standard API input, output, and cache-read rates are each half Opus 5.5's. The available evidence does not establish a decisive performance lead across the board, however, and long inputs are an exception to the price gap.

Key points

  • GPT-6.1 Sol succeeds GPT-6 Sol. It is not simply Astra sold more cheaply, but a separate model that changes the balance of capability and cost for complex work.
  • Both companies compete in code modification, computer use, and business workflows. Benchmark versions, reasoning effort, execution software, and model fallback conditions differ, so published results cannot be combined into one definitive ranking.
  • Sol's rates rise above 272,000 input tokens. Selection should account for total consumption, completion time, and human corrections as well as per-token prices.

Sources checked September 30, 2026. Announcement dates follow the official documents. This is an analysis of public information, not a LATENT review measuring both models under identical conditions.

Contents

  1. Why GPT-6.1 Sol belongs in a comparison with Opus 5.5
  2. Coding comparisons need matching tests and execution conditions
  3. Computer-use and business-workflow tests differ in tasks and scoring
  4. The API price gap narrows on long inputs
  5. Distinguish response speed from time to finish the work
  6. Availability and considerations when delegating actions to AI
  7. Choose for your own workload

Why GPT-6.1 Sol belongs in a comparison with Opus 5.5

Despite their similar names, GPT-6 Sol, GPT-6.1 Sol, and GPT-6 Astra should be distinguished. GPT-6 Sol and Luna became available September 22; GPT-6.1 Sol followed the next week. Astra remains a separate, higher-tier model. OpenAI's model documentation recommends comparing Sol with Astra on your own tasks to assess the quality and cost difference. Official changelog; GPT-6.1 Sol specifications

That positioning helps explain the competition with Opus 5.5. Anthropic presents Opus 5.5 as suited to extended agentic coding and knowledge work, and as a starting point for everyday workloads. The models overlap in tasks that involve using tools and completing multiple steps, rather than simply answering questions. Claude model overview

OpenAI's announcement includes cost-performance comparisons with Opus 5.5 on GDP.pdf, for understanding PDFs, and AutomationBench, for business workflows. The choice of comparisons shows that Opus is a competitive reference point. It does not establish development motives such as rushing the release to beat Opus.

If price alone is the criterion, Opus is not the only alternative. The current Claude Sonnet 5.5 also costs $2 per million standard input tokens and $10 per million output tokens, matching Sol. The useful question in comparing with Opus is where paying more still makes sense for complex work, not simply whether a cheaper AI exists. Anthropic pricing

Coding comparisons need matching tests and execution conditions

Software development demands more than generating a few lines of code. A model must understand an existing structure, edit multiple files, and continue until tests pass. DeepSWE evaluates extended development tasks of this kind. Its provider, Datacurve, says it creates tasks without copying existing patches and tests software behavior rather than the shape of the code. DeepSWE

According to OpenAI, GPT-6.1 Sol matched Astra on DeepSWE v1.1 at roughly one-fifth the cost. It also exceeded GPT-6 Sol's best score by 6.4 points with less reasoning effort and lower cost. In OpenAI's comparison, this creates more room to reduce the cost of work that previously required higher budgets or reasoning effort with Sol. GPT-6.1 Sol announcement

Opus 5.5's system card reports 74.2% on DeepSWE v1.1, averaged across five runs. However, the Datacurve public leaderboard retrieved for this article did not contain both models' new results, so we could not establish a direct comparison under matching execution conditions. System card, page 175; DeepSWE public leaderboard

Anthropic's announcement reports 66.4% on Terminal-Bench 4.0 and 57.8% on CursorBench 4.0, which includes tasks spanning multiple files. But it does not provide GPT-6.1 Sol results. Figures for GPT-5.6 Sol or Astra in a comparison published before Sol 6.1's release cannot be substituted for the new model's performance. Opus 5.5 performance table

The conditions also differ. Anthropic's table generally uses max reasoning effort, but its Terminal-Bench 4.0 result uses xhigh for Opus and OpenAI's published high result for Astra. Opus's 66.4% has a standard error of ±2.6 percentage points. According to the system card, Opus ran 66 tasks five times each in Claude Code's --bare mode, for 330 attempts. Safety-driven fallback to other models affected 10% of attempts. System card, page 178

The software that supplies tools and execution procedures also influences the result. In the FrontierCode evaluation published by Anthropic, for example, Claude ran in Claude Code while GPT ran in Codex CLI. A model name and a score alone conceal that difference. System card, pages 175–177

OpenAI also states that it evaluated its own models in research environments or through the API and took competitors' figures from public reports. That does not mean the prompts and tools matched everyday Codex or Claude Code use. Rather than declaring one model the coding winner from these sources alone, it is more reasonable to recognize the improvement over the previous Sol while assessing a direct match with Opus separately.

Computer-use and business-workflow tests differ in tasks and scoring

Computer use has improved as well. OpenAI says GPT-6.1 Sol scored seven points higher than GPT-6 Sol at maximum reasoning effort on OSWorld 2.0's offline tasks, at less than half the cost, and narrowed the gap to Astra to 2.1 points. These are scores including partial credit on the August 8, 2026 task set. OpenAI's computer-use evaluation

Anthropic's announcement reports 81.8%, including partial credit, for Opus 5.5 on OSWorld 2.1. Its system card labels the same 81.8% result 2.0 and says it used the September 10 task set. The naming discrepancy remains unresolved between the documents, but at a minimum, that task set differs from OpenAI's August 8 version. Announcement; system card, pages 205–206

Anthropic also changed from discarding earlier screenshots to retaining them and compacting the conversation after it exceeds 100,000 tokens. It specifies a 1080p Ubuntu environment, a maximum of 500 actions, max reasoning effort, and five independent runs. The system card cautions against directly comparing the result with earlier task sets or different execution setups.

Partial credit matters. Alongside Opus's 81.8% score, the share of tasks that satisfied every check was 48.7%. The 81.8% figure should not be read as the proportion of jobs fully completed, nor plotted beside Sol's result as if both were directly comparable success rates. System card, page 206

Some practical workloads do have published direct comparisons. On GDP.pdf, which asks expert questions about complex PDFs, OpenAI reports that GPT-6.1 Sol outperformed Opus 5.5 with model fallback across the reasoning settings tested, at less than half the cost per task. On AutomationBench, it says Sol at medium scored 2.2 points above Opus at roughly one-third the cost. OpenAI's practical-work evaluations

These are OpenAI's published comparisons, not LATENT measurements. The announcement text does not establish whether AutomationBench's 2.2-point difference is statistically meaningful. The Opus score in Anthropic's announcement was measured by Zapier during early access, without model fallback and with safety-triggered refusals counted as failures.

Work evaluated Published result reviewed Limits to the comparison
Understanding PDFs OpenAI reports Sol outperforming Opus Tests questions about PDFs, not the entire process of producing a deliverable
Multi-step business workflows OpenAI reports a score and cost advantage for Sol medium Requires matching benchmark versions, reasoning settings, and fallback conditions
On-screen actions Both companies report gains over their previous models Task-set dates and history management differ; a direct ranking is not justified

These distinctions also matter in daily use. Extracting a figure from a document and answering a question is different from using that figure to build a spreadsheet, enter data into an app, and finish the job. Strong reading performance alone does not establish reliable execution through to delivery.

The API price gap narrows on long inputs

The standard price difference is clear. The table below lists standard direct-API rates in US dollars per million tokens. Tokens are small processing units and do not correspond one-to-one with Japanese characters. These rates are also separate from monthly subscription prices and usage allowances. OpenAI pricing and specifications; Anthropic pricing

Billing category GPT-6.1 Sol Claude Opus 5.5
Standard input $2 $4
Cache reads $0.10 $0.20
Cache writes $2.50 $5 for 5-minute retention / $8 for 1-hour retention
Output $10 $20

These Sol rates apply to inputs of 272,000 tokens or fewer. Cache rates apply to the portion successfully reused. Faster processing, regional requirements, and tool fees have separate conditions.

Caching reuses computation for material or conversation prefixes supplied repeatedly. Extended development work often involves the same code and instructions, making cheaper cache reads useful. But the initial cache write also costs money. OpenAI's write rate is a separate rate applied to the relevant input portion, rather than a surcharge added to the normal input rate. OpenAI's caching guide

Retention conditions differ too. GPT-5.6 and later models, including GPT-6.1 Sol, have a minimum retention setting of 30 minutes from the last write or reuse. Anthropic offers 5-minute and 1-hour cache writes. Compare not just read prices but whether a cache will remain usable at your actual request intervals. OpenAI retention conditions; Anthropic cache pricing

Long inputs narrow the price gap. Once Sol's input exceeds 272,000 tokens, input and cache rates double and output rates rise by 50% for the entire request. That means $4 for standard input, $0.20 for cache reads, $5 for writes, and $15 for output. Opus 5.5, by contrast, applies its standard rates across its one-million-token context window. Sol's long-context conditions; Claude long-context pricing

Assuming identical billable token counts, a single uncached request would cost the following.

Assumed usage GPT-6.1 Sol Claude Opus 5.5
100,000 input + 10,000 output tokens $0.30 $0.60
300,000 input + 10,000 output tokens $1.35 $1.40

Rate-based estimates by LATENT. Output means all billable output tokens, including reasoning. The figures exclude cache writes and reads, tools, faster processing, regional requirements, and tax. Tokenization and consumption differ by model, so these are not measured costs for completing the same work.

In the second example, the difference is only five cents. Sending a large collection of documents in full every time produces a different cost advantage from extracting and reusing only what is needed.

Sol supports a 1.05-million-token context window and Opus one million; both have a maximum output of 128,000 tokens. Context is working space for input and generation, not a guarantee that every part of a document filling that space will be read with equal accuracy. Sol specifications; Claude specifications

Distinguish response speed from time to finish the work

Both models let users adjust reasoning effort. Sol supports low, medium, high, xhigh, and max, with medium as the default. Opus 5.5 also defaults to medium and offers settings from low to max. Sol does not support none or minimal, and Opus 5.5 cannot turn thinking off. Identical setting names are not a shared standard for computation or waiting time. Sol reasoning settings; Claude's effort guide; model overview

Less reasoning can make an individual attempt lighter. But if a difficult task then needs more trial and error, total completion time may increase. Maximum reasoning does not guarantee a higher score either. On Anthropic's FrontierCode v1.1 results, Opus medium scored 54.6%, compared with 54.4% for max in the performance table. Small differences like these should not be treated as conclusive evidence that one setting is better. Opus results by reasoning effort

Speed claims also need the right comparison. Anthropic's claim of more than 30% faster output compares Opus 5.5 with Opus 5, not Sol. It advertises up to 2.5x speed in a separately priced Fast mode. OpenAI announced that Sol's Ultrafast mode would offer up to 8x token generation speed relative to standard Codex speed in the following days. An announced future multiplier and current standard speed cannot establish a winner. The companies' announcements; Opus speed details

API prices checked September 30 list Sol Fast at $4 for input and $20 for output, and Opus Fast at $8 and $40. Opus Fast is a research preview with restrictions on where it is offered. This comparison also needs to be kept separate from standard pricing. OpenAI pricing; Anthropic pricing

Our document review did not establish a direct comparison of time to first response and time to completion with matching input lengths, reasoning settings, and execution environments. The speed users experience includes waiting for the first response, tool calls, and recovery from failures, as well as the rate at which text appears.

Availability and considerations when delegating actions to AI

At launch, GPT-6.1 Sol was announced for Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex, and for the API. OpenAI explicitly said it was not yet available in ordinary Chat. Actual access also depends on the plan, client, and organization settings. Its API identifier is gpt-6.1-sol, and tool-using workflows use the Responses API. OpenAI availability details; model specifications; changelog

Opus 5.5 is announced for Claude, Claude Code, and Claude Platform, as well as AWS, Google Cloud, and Microsoft Azure. Its direct-API identifier is claude-opus-5-5. Pricing and features through cloud providers should not be assumed to match the direct-API rates in this article. Anthropic availability details; models by platform

When delegating actions, the conditions under which a model stops are part of its performance. In Anthropic's evaluations, safety mechanisms can route cyber tasks to Opus 4.8 and tasks such as biology to Opus 5. Results that include fallback do not mean Opus 5.5 alone solved every task. They need to be distinguished from tests without fallback. Opus 5.5 evaluation conditions

OpenAI likewise says it applies Astra's safety measures to GPT-6.1 Sol. It evaluates attempts to bypass automated-review refusals and to perform actions the user did not authorize. These tests deliberately collect difficult situations; they do not measure the incident rate in ordinary use. GPT-6.1 Sol system card

Whichever model you choose, better instruction-following does not mean every action can be delegated without review. Deciding where a person checks difficult-to-reverse actions, such as sending email or deploying to production, belongs alongside model selection.

Choose for your own workload

GPT-6.1 Sol is worth comparing if you are considering Opus 5.5. Its lower standard and cache rates could be particularly useful for repeated code changes or document processing that reuse the same background material. Public evaluations also show improvements in practical capabilities that support that proposition.

The case for switching on price alone is weaker when long inputs narrow the gap, or when Opus finishes a task with fewer retries. This document comparison cannot establish which workloads fall into the latter category.

One approach is to choose examples from your own work—an existing-code fix, a report based on long documents, and a routine task requiring screen interaction—and apply the same acceptance criteria. Record completion time, human corrections, and reasons for stopping as well as billing details. Distinguishing work that is handled well at ordinary reasoning effort from work that benefits from more reasoning avoids using the highest-scoring configuration for every task.

This release does not settle a universal winner or loser against Opus 5.5. It adds a lower-cost option worth comparing when delegating complex work. The place to test that difference is the process of finishing your own tasks, not a single line in an overall ranking.