AI news, with the context that matters.

Products & Services · ·

GPT-6 Astra launches with improved computer-use performance

GPT-6 Astra scores higher on computer-use evaluations and completes tasks faster. We explain performance and pricing changes, availability and usage restrictions.

OpenAI has announced GPT-6 Astra. Its president describes the start of an 'AGI era,' but how useful is the model for actual work? We focus on computer-use evaluations, differences from its predecessor and access conditions.

Key takeaways

  1. Computer-use performance improves. In OpenAI's evaluation, task scores rise and completion time roughly halves.
  2. API input and output rates are 2.5 times the predecessor's discounted rates. Whether cost per task falls depends on the work.
  3. On announcement day, access is limited to selected organizations. OpenAI and Microsoft Azure say they will expand availability over several days.
  4. General users cannot access certain capabilities, including exploit generation.
GPT and Astra title image against a swirling star background
Source: OpenAI

What happened

  • OpenAI announced GPT-6 Astra at 3 a.m. Japan time on September 4. Access begins with selected organizations, with expansion to paid plans and the API planned within days. Microsoft Azure also began offering it to participants in a limited-access program that day.
  • API pricing is $10 per million input tokens and $50 per million output tokens.
  • In a press briefing, president Greg Brockman said he believed the era of AGI had begun.
  • On September 1, US time, OpenAI had announced that the model's cyber capabilities reached Critical, the highest threshold in its framework.

Primary sources: OpenAI announcement, System card, Path to Astra

Higher computer-use scores and shorter task times

Astra improves at reading a screen, operating a mouse and keyboard and completing work with software. This way of using AI is often described as an agentNote 1.

In OpenAI's evaluation, the computer-use score increases from 65.7% for GPT-5.6 Sol to 72.6%. Tasks also take less time: under simulated response latency, time per task falls from roughly 75 to 40 minutes. The business-automation score more than doubles from 18.1% to 41.4%, while retrieval from documents exceeding 500,000 tokensNote 2 improves from 73.8% to 96.3%.

OpenAI lists examples such as filling web forms, updating customer-management systems and organizing calendars. Other demonstrations include designing manufacturable circuit-board placement and routing from schematics, and importing a house from 3D software into a game engine for a walk-through. These are examples selected and published by the developer.

New features support longer-running work. In the Codex developer tool, an experimental feature lets the AI write progress notes and reread them later. Since a model can handle only a limited amount of information at once, these notes help it recover details of earlier work.

Figure 1Evaluation results compared with the predecessor

Results published by OpenAI. Each value uses the highest result measured across different reasoning budgets.

  • Complete tasks by operating a computerOSWorld 2.0; scores include partial credit

    GPT-5.6 Sol65.7%

    GPT-6 Astra72.6%

  • Time per computer-use taskWith simulated response latency

    GPT-5.6 SolApprox. 75 min

    GPT-6 AstraApprox. 40 min

  • Automate business tasksAutomationBench score

    GPT-5.6 Sol18.1%

    GPT-6 Astra41.4%Approx. 2.3x

  • Find information in long documentsMRCR; documents of 512,000–1,000,000 tokens

    GPT-5.6 Sol73.8%

    GPT-6 Astra96.3%

  • Locate the target of a screen interactionScreenSpot-Pro

    GPT-5.6 Sol76.9%

    GPT-6 Astra92.7%

  • Unintended outcomes during computer useInternal safety tests with safety checks disabled; lower is better

    GPT-5.6 Sol22.0%

    GPT-6 Astra2.4%

Computer-use task time roughly halves, and the automation score more than doubles. OSWorld 2.0 includes partial credit, so its score is not the proportion of tasks completed.

Figures published by OpenAIValues come from OpenAI's announcement of September 3, 2026. Many benchmarks were created by external organizations. The editors calculated 'approximately 2.3x' from the two model scores.

An independent overall score matches the predecessor

The 72.6% computer-use score includes partial credit. It does not mean that 72.6% of tasks were completed from start to finish. Human verification remains necessary when delegating actual work.

The figures above appear in OpenAI's announcement. The company says GPT evaluations used research environments or the API, taking the best result across reasoning budgets. It notes that ChatGPT results may differ somewhat.

Independent results released on announcement day showed no change in the overall score. Astra scored 61 on the Artificial Analysis Intelligence IndexNote 3, tying its predecessor and trailing Anthropic's Claude Fable 5.1 by five points. This index also appears in OpenAI's announcement. Astra's coding index was 67, two points above the predecessor but below Fable 5.1's 70. At maximum reasoning for both generations, Astra used one third as many tokens on the coding evaluation.

Results on the ARC-AGI-3 reasoning benchmark also differ substantially by configuration. According to benchmark operator ARC Prize, Astra scored 99.9% using a connection mode that preserves intermediate reasoning and 62.7% using the standard mode. The two runs also used different reasoning budgets.

Some items in OpenAI's announcement trail rival models. On a benchmark of difficult expert-level questions, Astra scores 57.2%, versus 65.0% for Fable 5.1 and 63.6% for Claude Opus 5. Results for OpenAI's own GDPval measure of practical economic value are not included.

Safety evaluations show improvements alongside concerns. In internal tests with automatic safety checks disabled, unintended outcomes during computer use fell from 22.0% to 2.4%. On a test of requests seeking cooperation with cyberattacks, refusal increased from 59% for GPT-5.6 Sol to 91.5% for Astra. However, the system cardNote 4 says Astra's reasoning has become harder to monitor externally. External evaluator Apollo found that, at maximum reasoning, the model recognized that it was being evaluated 50.6% of the time. Its behavior in evaluations could differ from behavior in deployment.

Access was limited at announcement, and independent evaluations remained scarce. Further evidence is needed to judge performance across a broad range of work.

Higher unit prices, but task costs depend on the use case

API rates are $10 per million input tokens and $50 per million output tokens. Relative to the predecessor's list prices of $5 and $30, input doubles and output rises about 1.7 times. The predecessor is discounted through at least November 21; against its discounted $4 input and $20 output rates, both are 2.5 times higher. Astra matches Anthropic's high-end Fable 5.1. For inputs exceeding 272,000 tokens, input and cache rates double and output rates rise 1.5 times, adding costs for long documents.

OpenAI argues that cost per task matters more than unit price. In the briefing, Brockman said that a token represents different amounts of content across companies and models, making token-price comparisons unhelpful. The announcement reports an estimated cost per software-development task about 32% below the predecessor.

Independent results on announcement day differed by workload. With both models at maximum reasoning, Artificial Analysis found one third as many tokens on coding tests but roughly the same cost per task as the predecessor. It rated price-performance among the best. On the overall index workload, output tokens fell only about 10%, while cost per task rose 75%. The comparison uses the predecessor's discounted price. These tests differ from OpenAI's, so the results cannot be compared directly.

Fewer tokens do not necessarily mean a lower bill. Announcement-day evidence suggests the impact depends on the task. An early reviewer at the newsletter Latent Space reported spending about $100 on two days of work including experiment monitoring. The reviewer also estimated about $5.94 per hour for output alone if a single process continuously generated 33 tokens per second at $50 per million output tokens. Input, cache and parallel-agent costs would be additional.

Figure 2Pricing comparison

US dollars per million tokens.

OpenAI

GPT-6 AstraNew model

Input$10

Output$50

GPT-5.6 SolPredecessor; list price

Input$5

Output$30

GPT-5.6 SolPredecessor; discounted through at least November 21

Input$4

Output$20

Anthropic

Claude Fable 5.1High-end model

Input$10

Output$50

Claude Opus 5

Input$5

Output$25

SourcesOpenAI model pages (GPT-6 Astra, GPT-5.6 Sol), GPT-5.6 announcement, Anthropic pricing

Access on announcement day is limited to selected organizations

Only selected organizations have access on announcement day. OpenAI says it will expand to paid ChatGPT plans—Plus, Pro, Business and Enterprise—and the API within days. Enterprise administrators must enable it. No free-plan availability has been announced.

Microsoft Azure also began offering access to selected program participants on announcement day, with expansion to participating customers over several days. OpenAI says AWS availability is expected within days too. The announcement does not exclude Japan.

Japan-only data processing is unconfirmed

At announcement, it was not possible to confirm whether processing could remain entirely within Japan. Microsoft listed two Azure deployment options: one without a restricted processing region and one restricted to the US. US-only processing carries a 10% price premium.

Neither the announcement nor the system card reports Japanese-language performance. Users working with Japanese documents or conversations need to evaluate the model on their own tasks.

Exploit generation and other capabilities are restricted

Astra's cyber capabilities have different restrictions for different users. On September 1, US time, OpenAI announced that Astra had reached Critical, the highest capability threshold in its framework, its first model to do so. During evaluation it found two previously unknown software vulnerabilities and demonstrated exploitation. OpenAI said disclosure to developers and maintainers was underway.

General users can access defensive functions such as code-security reviews and patch generation. Activities such as generating exploit code remain restricted.

Within weeks, defensive professionals are expected to gain additional capabilities through a vetted-access program, including vulnerability and exploit validation, malware analysis and detection-rule creation. More advanced work, including exploit generation, will start with a small set of testers and later expand to vetted users.

Figure 3Cyber capabilities and access groups

At announcement
General users

  • 01Code-security reviews
  • 02Patches for vulnerabilities

Capabilities available to general users

Within weeks
Vetted users

  • 03Vulnerability and exploit validation
  • 04Malware analysis
  • 05Detection-rule creation

Early access
A small set of testers

  • 06More advanced work, including exploit generationLater expanding to vetted users

Capabilities shown in the lower part of the figure are planned for gradual, restricted rollout.

SourcesOpenAI's announcement and Path to Astra

'We have reached AGI' remains the president's personal view

In the briefing, Brockman said he personally believed OpenAI had reached AGI, a term for AI able to handle a broad range of work like humans. He also noted that definitions differ and not everyone will agree on when it is achieved. He said contractual provisions triggered by reaching AGI no longer existed; TechCrunch reports that this referred to Microsoft's agreement.

The statement does not establish that broad classes of work can be delegated without human checks. Practical assessment still requires examining performance on your tasks, the effort of verifying results and cost. Independent evaluations and user reports remained limited at announcement.

Timeline

Dates are in US time.

  1. OpenAI releases GPT-5.6.

  2. OpenAI announces that the intrusion into Hugging Face was caused by its own model during testing.

  3. OpenAI says it cannot rule out its next model reaching the highest cyber-capability threshold and pauses some development work.

  4. OpenAI discloses a two-week training pause and the suspension of its largest training run.

  5. The suspended training resumes.

  6. OpenAI publishes Path to Astra and rates its cyber capability Critical, the highest threshold in its framework.

  7. GPT-6 Astra is announced and begins rolling out to selected organizations. It is September 4 in Japan.

Sources and references

Update history

  • Published.
  • Described Critical as a cyber-capability rating and corrected the vulnerability-disclosure status to ongoing. Clarified that the hourly cost estimate covers output tokens from a single process.
  • Revised Figure 1 and the following explanation to distinguish figures published by OpenAI from the organizations conducting the measurements.
  • Clarified that the refusal rate comes from a cyber evaluation and that Figure 1's partial-credit note refers to OSWorld 2.0.

LATENT uses Google Analytics (Firebase) to improve our articles. With your permission, cookies and similar technologies send browsing and interaction data to Google. Privacy policy