AI news, with the context that matters.

Columns · ·

Can efficiency gains offset Codex Pro's reduced usage allowance?

The new $200 Pro multiplier is 10x Plus. We examine actual usable capacity and efficiency gains in light of Tibo's additional explanation: a higher usage baseline, temporary retention of 20x for existing subscribers, and extra credits.

The $200 Pro plan's multiplier is changing from 20x Plus to 10x. That does not establish that every user immediately gets half as much usage. Tibo Sottiaux, who leads Codex and ChatGPT Work, also described a higher usage baseline and transitional arrangements for existing subscribers.

In his initial September 29 post, Japan time, Tibo said the API-equivalent allowance would halve. A follow-up at 00:10 Japan time on September 30 described a higher baseline including Plus, the new multipliers, and temporary retention of 20x plus extra credits for existing $200 Pro subscribers. Both explanations need to be considered when assessing actual usable capacity.

The key questions are how much the baseline rises, how long the transitional arrangements last, and how much less allowance the same work consumes. In addition, the initial post compares with one month ago, rather than immediately before the change, and Sol and Luna's price reductions cannot simply be applied to Astra.

Updated September 30, 2026. Based on the follow-up post, we have corrected the original article's description of the move to 10x for $200 Pro and the treatment of existing subscribers as unconfirmed. The opening, calculation assumptions, and conclusion have also been revised.

Key points

  • The announced new multipliers are 1x for Plus, 5x for $100 Pro, and 10x for $200 Pro. Because Tibo also says the baseline is rising, the change from 20x to 10x does not by itself establish a halving of the absolute allowance.
  • Existing $200 Pro subscribers are to retain 20x temporarily and receive additional credits. The post does not specify how long this lasts or how many credits they receive.
  • If the API-equivalent allowance actually halves, the cost per task must also halve to preserve the same amount of work at the same quality. This is a conditional calculation, not a measurement of every subscription after the transition.
  • The reported Astra improvements apply to some particularly high-consumption cases. The sources reviewed here do not show that every user's work now consumes half as much allowance or less.

Contents

  1. The usage multiplier is changing from 20x to 10x
  2. Tibo's baseline is Pro one month ago
  3. A 20% consumption reduction cannot offset a halved allowance
  4. Astra has not received the same price cut as Sol and Luna
  5. Caching can lower costs without reducing token counts
  6. Existing $200 Pro subscribers temporarily retain 20x
  7. Compare the allowance consumed to complete the same work

The usage multiplier is changing from 20x to 10x

In the initial post, Tibo said that reopening new $200 Pro subscriptions would come with a change in how usage is calculated, leaving half the API-dollar equivalent of the old plan. His wording was “half the dollar in API spend.” Initial post on X

In the follow-up, he said that Plus and Pro would be able to do more and that compute had been brought online to support the increase. As the baseline rises, he said, the relative multipliers between plans would change as follows. Follow-up post on X

Plan Announced new multiplier
Plus 1x
$100 Pro 5x
$200 Pro 10x

Multipliers are relative to Plus. The temporary retention of 20x and extra credits for existing $200 Pro subscribers are transitional arrangements discussed below. The post describes the policy; it does not establish that rollout has completed for every account.

Halving a relative multiplier and halving the absolute allowance are different things. If, on a comparable scale, the Plus baseline allowance doubled, 10x the new baseline would equal 20x the old one. If the baseline rose by 50%, it would amount to 75% of the old $200 Pro allowance. These examples illustrate the relationship between multipliers; the post does not specify the actual increase in the baseline allowance.

The follow-up does not explicitly retract the initial statement about halving the API-equivalent allowance. Nor does it quantify how the old and new baseline allowances map to API dollars. These two posts alone establish neither that everyone's absolute allowance halves because the multiplier halves, nor that everyone necessarily gets at least as much as before because the baseline rises.

Three things need to be distinguished: tokens, the units used to process text; what those tokens would cost through the API; and the allowance included in a subscription. When prices fall, the same number of tokens has a lower API-dollar value. More expensive models or processing settings can have the opposite effect.

Half the API-dollar value therefore cannot be read as half the tokens across every model. However, if the available API-equivalent allowance really halves while prices and the way tasks are carried out stay unchanged, the amount of work completed falls. The follow-up makes it even more necessary to distinguish which subscriptions this condition applies to, and when.

The full post also promises not to reintroduce Pro's five-hour limit and previews additional features that will not consume the allowance. The former lets users spend their weekly allowance when they choose; it does not increase the weekly total. The value of additional features also needs to be assessed separately from the amount of coding usage available. Archived full text

Tibo's baseline is Pro one month ago

An easily missed detail is that Tibo compares the amount of work users can finish with what they could do “one month ago.” Archived full text

OpenAI announced GPT-6 Sol and Luna on September 22, with API prices below the previous generation's promotional rates. The cheaper models therefore arrived before this explanation of the allowance change. OpenAI's announcement

Someone using the old model a month ago and someone already using the new model immediately before the change have different baselines. The first comparison can include the benefits of a model upgrade. The second cannot count the same price reduction again as a new benefit of the allowance change.

Suppose a task costs $1 with the old model and $0.50 with the new one. If the allowance falls from $100 to $50 in API-equivalent value, the old-model comparison is 100 tasks versus 100 tasks. But if the user could already use the new model with the $100 allowance before the change, the comparison becomes 200 tasks versus 100.

This is an illustration of different baselines, not a claim that the actual allowance was $100. It does not incorporate the higher baseline or transitional arrangements described in the follow-up. Without the actual old and new allowances, we cannot conclude either that usage halves as in this example or that being better than one month ago means there is no loss relative to immediately before the change.

A 20% consumption reduction cannot offset a halved allowance

For comparable tasks completed to the same quality, the arithmetic is straightforward.

Work completed = available API-equivalent allowance ÷ API-equivalent cost per completed task

The following table is a conditional calculation that applies only if the available API-equivalent allowance actually halves. It does not establish what fraction of the old allowance the new 10x plan provides, or when this condition would apply to existing subscribers temporarily retaining 20x.

Change in cost per task Work completed relative to before
No cost reduction 50% of previous amount
20% lower cost 62.5% of previous amount
50% lower cost 100% of previous amount
60% lower cost 125% of previous amount

Illustrative calculations by LATENT. Each row uses 0.5 divided by the new-to-old cost ratio per task. Task scope and quality are held constant, and costs, including retries, are assumed to change as shown. These are not measured Pro task counts or a reconstruction of its unpublished usage formula.

“Consumption also falls” describes a direction, not proof that the improvement is large enough. A 20% saving is welcome, but with half the allowance, the amount of work completed falls by 37.5%. Offsetting the allowance cut requires a 50% reduction in cost per task; doing more work requires a larger reduction.

Cost here means the API-equivalent cost of the entire task, accounting for price changes, caching, reasoning, and retries. Shorter output or faster responses alone do not establish that this condition is met.

Astra has not received the same price cut as Sol and Luna

The September 22 announcement and official model specifications checked on September 30 show the following standard API rates. Prices are in US dollars per million tokens; input prices are for uncached input.

Model comparison Input Output
GPT-5.6 Sol → GPT-6 Sol $4 → $2 $20 → $10
GPT-5.6 Luna → GPT-6 Luna $0.20 → $0.10 $1.20 → $0.50
GPT-6 Astra at the time of checking $10 $50

Sources: the Sol and Luna announcement, and official specifications for Sol, Luna, and Astra. Previous-generation figures are the promotional prices listed in the announcement. This table excludes caching, long-context, and fast-processing conditions.

For Sol, unchanged input and output volumes would make this portion of the cost half as much. If the allowance also halves, usage stays at the previous level. Any improvement beyond that would come from the model completing work in fewer steps.

Luna's output price falls from $1.20 to $0.50, a reduction of approximately 58.3%, rather than exactly 50%. In a simplified output-only comparison with unchanged token counts, half the allowance would still buy 1.2 times the old model's output volume. Real tasks include input and caching, so this is not a prediction of 20% more work in every case.

Astra cannot be assumed to receive the same halving of token prices shown for Sol and Luna. For people continuing to use Astra, the reduction in consumption per task needs separate examination. Saving money by switching to Sol and continuing to use Astra as before are different benefits.

There have, however, been reports of improvements to Astra. On September 6, Tibo said that changes for power users accessing Astra through a ChatGPT account could reduce the allowance consumed in some high-consumption cases to as little as one-third or one-quarter of the previous amount, without changing quality.

Where those improvements apply, they could more than offset a halved allowance. But these are maximum gains in some cases, not an average for all users. The improvements also predate the latest announcement, so they are not established as an additional benefit relative to immediately before the change.

Caching can lower costs without reducing token counts

Caching is one concrete mechanism behind lower consumption costs. It reuses processing of input shared with earlier requests, reducing the need to process the same context from scratch. OpenAI says it improved caching for GPT-6 and prices eligible cached input reads 90% below ordinary input. Official caching announcement

The GitHub example in that announcement reports a reduction of more than 50% in the share of input tokens requiring fresh processing, across billions of requests. It does not report a halving of total tokens, output, reasoning, and overall task costs. Nor can a Copilot result be applied directly to every Codex Pro user.

Sending the same million tokens can cost less if more of them qualify for cheaper cached reads. Conversely, tasks that introduce large amounts of new material or depend heavily on output may not see their total cost halve through input caching alone. Saying that token consumption also falls obscures this distinction.

Existing $200 Pro subscribers temporarily retain 20x

The follow-up says existing $200 Pro subscribers will keep the 20x multiplier for a while and receive additional credits. This establishes the difference in treatment between new and existing subscribers that the original article described as unconfirmed. It should be understood as a transitional arrangement with an unspecified end date, rather than permanent grandfathering. Follow-up post on X

The post does not specify the duration of the 20x extension, the amount, delivery date or expiry of additional credits, or the increase in the baseline allowance. Usage supported by extra credits during the transition should be assessed separately from the recurring weekly allowance afterward.

When checked for the original article, the official pricing page still described 5x and 20x, while the Pro help page described the pause on new subscriptions and its effect on existing subscriptions. Those pages alone are not grounds to dismiss the move to 10x or the transitional arrangements in the follow-up. We distinguish the announced policy from the figures and deadlines that remain unspecified.

A higher baseline, additional credits, and new features can each add value to the subscription. But people who already use their full coding allowance need to know whether they can sustain the same amount of work after the transition ends. That requires knowing how the recurring weekly allowance relates to consumption per completed task.

Compare the allowance consumed to complete the same work

The useful measure is not a large aggregate token count, but how many percentage points of the weekly allowance are consumed to finish the same work at the same quality. If a comparable fix used 2% of the weekly allowance before the change and 4% afterward, the number of such fixes that fit within the allowance halves, all else being equal. A real comparison should cover multiple tasks and retain the variation between them.

Three sets of records are enough to frame the comparison.

What to record What it helps establish
Completion criteria and number of retries Whether savings come at the expense of quality
Model, reasoning and speed settings, and the input, cached-input, and output breakdown Why the cost changed
Applicable multiplier and transitional arrangements; weekly allowance usage and resets, separating additional credits How many comparable tasks fit during and after the transition

Periods containing other concurrent work or an allowance reset cannot be compared directly. Record whether the account retains 20x, uses the new 10x allowance, or is consuming additional credits. If reporting API-equivalent costs, distinguish a comparison repriced using one common rate card from one using the rates in effect at each point in time. This separates allowance and transition effects from processing efficiency and price reductions.

This investigation did not measure matched before-and-after tasks under those conditions. The follow-up establishes an announced policy of 10x for the new $200 Pro plan, temporary retention of 20x plus extra credits for existing subscribers, and a higher baseline. It does not yet establish the change in actual usable capacity, or that everyone, including people who mainly use Astra, can do at least as much as before after the transition ends.

Evaluating Tibo's explanation requires accounting for the higher baseline and transitional arrangements as well as model efficiency. How much of the work users could complete immediately before the change can they still do at the same quality and monthly price, during and after the transition? That calls for measurement, rather than a conclusion based only on the multipliers.

Sources and scope

Sources for the original article were checked between midnight and 1 a.m. Japan time on September 30, 2026. This same-day update corrects the omission of the follow-up posted at 00:10. Statements, official prices, and LATENT's illustrative calculations are distinguished throughout; announced policy is not treated as confirmation of completed account rollout. The API price table reflects the original checking time.

  1. Tibo's post about the $200 Pro plan — September 29, 2026, 06:41 UTC. The opening and timestamp were confirmed through X's official embed feeds. Because those feeds truncate the long-form post, the remainder was cross-checked against recodex's full-text record and a Reddit reproduction. The reopening and additional features are treated as announced plans, not confirmed completed rollouts.
  2. Tibo's follow-up explanation of $200 Pro — September 29, 2026, 15:10:31 UTC, or September 30 at 00:10:31 Japan time. The timestamp and opening were confirmed through X's official embed feed. Because that feed truncates the long-form post, the full text was checked through FxTwitter's retrieval. The later section, including the move to 10x and transitional arrangements for existing subscribers, was therefore verified through a third party.
  3. Introducing GPT-6 Sol and Luna — OpenAI, September 22, 2026. New and previous model API rates and reported improvements.
  4. GPT-6 Sol, GPT-6 Luna, and GPT-6 Astra — OpenAI model specifications, including standard rates at the time of checking.
  5. Tibo's post about Astra usage improvements — September 6, 2026. The full post was confirmed through X's official embed feed. Its limited scope and maximum reported improvement are preserved.
  6. Better prompt caching for GPT-6 — OpenAI, September 22, 2026. Caching improvements and the GitHub report reproduced by OpenAI.
  7. ChatGPT Work and Codex pricing, and About ChatGPT Pro tiers — Official explanations as checked for the original article, not treated as updated terms incorporating the follow-up post.