AI news, with the context that matters.

Products & Services · ·

GPT-6 Sol and Luna launch with API rates cut by half or more

In independent evaluation, cost per task falls about 47% for Sol and 61% for Luna. Overall performance is near the previous generation, with gains in avoiding wrong answers and business automation but weaker results in document creation and other work.

For long AI tasks, the cost of finishing matters more than the price of a single answer. How much do Sol and Luna's price cuts help? We compare predecessor pricing and independently measured work quality.

Key takeaways

  1. API input and output rates fall by half or more, and independent evaluation also finds lower cost per task.
  2. Overall performance is near the previous generation. Avoiding wrong answers and business automation improve, while document creation and some other areas decline.
  3. At announcement, the models are available in ChatGPT Work, Codex and the API, but not ordinary Chat.
Official GPT-6 Sol and Luna announcement image
Source: OpenAI

What happened

  • OpenAI announced GPT-6 Sol and Luna on September 22, at 3 a.m. September 23 in Japan.
  • Rollout begins for eligible ChatGPT Work and Codex plans, with API access also available.
  • Standard API rates per million tokens are $2 input and $10 output for Sol, and $0.10 input and $0.50 output for Luna.

Primary sources: OpenAI announcement, Sol and Luna specifications

Sol targets complex work; Luna targets inexpensive high-volume processing

GPT-6 Sol and Luna follow Astra, released earlier in the month. OpenAI positions Sol for complex programming and multi-step work, and Luna for high-volume, narrowly scoped processing. It continues to recommend Astra for the most demanding work.

On that basis, Sol is a candidate for investigating multiple files and changing a program, while Luna is a candidate for classifying inquiries or extracting fields from documents at scale. These are guidelines, not rules: cheaper models can meet the required accuracy on some tasks, and difficult work need not automatically go to Sol.

Both API models support up to 1.05 million tokensNote 1 across input and output, with a maximum of 128,000 output tokens per response. They accept text and images and produce text. Equal capacity does not establish equal accuracy on details in long documents.

API input and output rates fall to half or less

A major change is lower API pricing. Per million tokens, Sol input falls from $4 to $2 and Luna input from $0.20 to $0.10. Sol output falls from $20 to $10 and Luna output from $1.20 to $0.50.

Sol's input and output rates both fall 50%. Luna's input falls 50% and output about 58%. The comparison is with already-reduced GPT-5.6 Sol and Luna prices. OpenAI calls both comparison prices promotional, but Sol's rate has a stated period through at least November 21, while Luna's is the rate reduced on July 30. No end date for Luna's reduction is given in that announcement. These are not comparisons with original launch prices.

Standard input and output rates for the new Luna are one twentieth of Sol's. But token usage and retries differ by model, so the cost of completing a task is not necessarily one twentieth.

Figure 1API input and output pricing

US dollars per million tokens, comparing standard processing with at most 272,000 input tokens.

  • Sol input50% reduction

    GPT-5.6$4.00

    GPT-6$2.00

  • Sol output50% reduction

    GPT-5.6$20.00

    GPT-6$10.00

  • Luna input50% reduction

    GPT-5.6$0.20

    GPT-6$0.10

  • Luna outputAbout 58% reduction

    GPT-5.6$1.20

    GPT-6$0.50

Predecessor rates are their reduced prices. Cache reads and writes and long-input surcharges are excluded.

Official pricingBased on OpenAI's announcement and Sol and Luna specifications. Reductions were calculated from published rates.

Independent evaluation also finds lower cost per task

In its announcement-day evaluation, Artificial Analysis published cost per task for the Intelligence Index. Sol falls from $1.99 to $1.06, and Luna from $0.18 to $0.07. Calculated from the published figures, those are reductions of approximately 47% and 61%.

Artificial Analysis describes overall performance as broadly comparable to GPT-5.6. The striking result is therefore lower cost for roughly the same overall level of work, rather than a large jump in aggregate performance.

The savings do not come from lower usage. Output per task increases from roughly 29,000 to 31,000 tokens for Sol and from about 41,000 to 51,000 for Luna. Since cost falls despite that increase, the evaluator attributes the savings to lower prices.

All models were measured at maximum reasoning, which differs from API defaults. A task here is an evaluation task, not the average cost of an inquiry or code change.

Figure 2Cost per task in independent evaluation

Measured on the Artificial Analysis Intelligence Index, with maximum reasoning for every model.

  • SolAbout 47% lower

    GPT-5.6$1.99

    GPT-6$1.06

  • LunaAbout 61% lower

    GPT-5.6$0.18

    GPT-6$0.07

Overall performance is near the previous generation. This compares the cost of tokens consumed by evaluation tasks, not API unit rates.

Independent measurementsBased on Artificial Analysis's September 22 evaluation. Published costs are used to calculate reductions. These figures are not the cost per real-world business task.

Some capabilities improve while others trail the predecessor

In Artificial Analysis's programming evaluation, the models diverge. On the Coding Agent Index measured through Codex, Sol rises from 55 to 57 while Luna falls from 43 to 41. Both comparisons use maximum reasoning.

On the AA-Omniscience knowledge test, the 'hallucination rate' falls from 92% to 60% for Sol and from 93% to 77% for Luna. This metric is the share of wrong answers among wrong answers, partially correct answers and abstentions. Sol abstains more often and produces fewer wrong answers, but its accuracy across all questions falls from 59% to 54%. Luna's accuracy is nearly unchanged, from 43% to 44%.

Both improve on AutomationBench-AA business automation and Terminal-Bench 4.0 command-line work. Both decline on GDPval-AA v2.1, which covers work across 44 occupations. Visual inspection by the evaluator found that weaker presentation and missing required elements tended to depress scores.

OpenAI also published results showing gains in programming and computer use. It says its own models were tested in research environments or the API and may differ slightly from ChatGPT output. Competitor results were taken from published reports. This is therefore not a retest of every model under one shared environment.

When considering a switch, compare your actual tasks alongside prices and aggregate scores. For document creation, check appearance, required sections, sources and numbers. More human revision can replace the time and money saved on API calls.

Prices rise above 272,000 input tokens

The pricing rules for Sol and Luna include a long-input surcharge. Above 272,000 input tokens, input rates double and output rates rise 1.5 times for the entire request, not only the excess portion.

Rates then become $4 input and $15 output for Sol, and $0.20 input and $0.75 output for Luna, per million tokens. Supporting 1.05 million tokens does not mean the whole capacity is available at standard rates.

For Japanese prose, 272,000 tokens is roughly 350,000–420,000 characters, or around 900–1,050 sheets of 400-character manuscript paper in an illustrative calculation. This is not measured Sol or Luna usage. For English, OpenAI's rule of thumb of about four characters per token gives approximately 1.09 million characters.Note 1

Image usage depends on the model and detail setting as well as resolution. As a reference example, a 1,024-by-1,024 image at high detail in GPT-6 Astra uses 1,229 tokens in an official calculation. Images alone would approach 272,000 tokens at roughly 220 images. That document does not give Sol or Luna image accounting, so the same image count cannot serve as a limit for either model.

The 272,000-token threshold applies to the whole input, not just the newly typed question. Documents, images, conversation history and application instructions count too. A short question can approach the threshold when accompanied by substantial material or history.

Reusing repeated instructions and documents from cache can reduce cost. Cache reads are 90% cheaper than ordinary input, but writes cost 1.25 times the input rate. The benefit differs between one-off documents and frequently reused material.

Estimating cost requires considering input and output volume, long-input surcharges and the fraction that can be reused from cache. Unit-price cuts alone do not establish that the monthly bill will halve.

The models are not yet available in ordinary Chat

At announcement, availability covers ChatGPT Work, Codex and the API. Work and Codex roll out to Plus, Pro, Business, Enterprise and Edu users. Free and Go users can use Luna in the desktop app.

Neither model is available in ordinary Chat at announcement. ChatGPT rollout is staged over the day, so not every user's interface changes immediately.

API model names are gpt-6-sol and gpt-6-luna. The reduction to half or less concerns API input and output rates, not an announcement that ChatGPT subscriptions will halve in price.

Timeline

Dates follow the original announcements. This launch occurred at 3 a.m. September 23 in Japan.

  1. OpenAI announces the previous GPT-5.6 series.

  2. GPT-6 Astra is announced and begins rolling out to selected organizations.

  3. GPT-6 Sol and Luna are announced with lower API rates than their predecessors.

Sources and references

Update history

  • Published.
  • Added illustrative token counts in characters and images.
  • Expanded on improvements in independent evaluation and the definition of hallucination rate, updating the description and takeaways. Corrected predecessor pricing conditions and evaluation explanations, and changed the archived link for the image-accounting example.

LATENT uses Google Analytics (Firebase) to improve our articles. With your permission, cookies and similar technologies send browsing and interaction data to Google. Privacy policy