Independent research · 13 September 2026 · USD

The real economics
of an AI subscription.

From a $200 Codex plan to billion-dollar compute fleets: a ten-developer comparison of token value, customer margins, weekly demand and operating costs.

Published prices & limitsExplicit financial scenariosNo private telemetry
Astra vs Sol token cost
2.5×

Published short-context API and Codex credit-weight ratio. [1, 2]

Fixed weekly Pro tokens
Unknown

No public, guaranteed weekly token allocation. Five-hour estimates are not a weekly cap.

OpenAI operating-cost model
$6B/mo

Illustrative central case; $4–8B range. Not a disclosed company result.

Z.ai Max calculable example
389–779M

Weekly GLM-5.3 tokens at a 15/80/5 mix; peak versus off-peak. [11, 12]

What can actually be concluded: under the same weighted credit budget and token mix, Sol permits roughly 2.5× Astra's token volume, while the API-equivalent dollar value stays similar. Public information does not establish the absolute weekly budget of a $200 Pro account.

Scope: OpenAI, Anthropic, Google, xAI, Meta, DeepSeek, Alibaba, ByteDance, Moonshot and Z.ai. This is a selected group of major model developers, not a definitive valuation or intelligence ranking. All financial estimates are independent scenarios. USD prices exclude taxes, app-store differences and negotiated contracts.

01 / Your $200 Codex question

A reproducible conversion, with the unpublished variable left visible rather than invented.

Official limits

Five-hour windows ≠ weekly tokens

Pro 20× documentation estimates 100–900 local Astra messages or 200–2,000 local Sol messages per five-hour window. Weekly limits also apply. Cloud chats use Sol. Message cost varies with work and context. [2]

Astra credits / 1M tokens: 250 input · 25 cache · 1,250 output
Sol: 100 input · 10 cache · 500 output

At the default mix: Astra = 120 credits/M; Sol = 48 credits/M. If the actual included weekly balance is C credits, tokens in millions are C/120 and C/48. C is not published.

Sensitivity calculator

Output fixed at 5%; fresh input is the remainder. The $300 default is an illustration, not an estimate of your entitlement.

Astra tokens/week
Sol tokens/week
Assumed weekly API valueAstra tokensSol tokensMonthly API valueValue / $200 fee
$15031.25M78.13M$6503.25×
$30062.50M156.25M$1,3006.50×
$600125.00M312.50M$2,60013.00×

These three budgets are sensitivity points, not a probability range. Month = 52/12 weeks. “Tokens” includes repeated cached context and billable reasoning, not only visible output. Fast mode, long context, tools, cache creation and storage can change cost. If all 15% fresh input incurs 1.25× cache-write pricing, blended rates rise from $4.80/$1.92 to $5.175/$2.07 per million.

02 / What an identical token workload costs

Representative text models. Price per million tokens; standard short-context processing where applicable. Not a quality-equivalence ranking.

Last two columns recalculate.

Swipe horizontally to inspect the full table.

Developer / modelFresh / MCache / MOutput / MCost / 100M$200 API buys

A dash means a rate was not independently verified here, not that caching is free or unavailable. Meta uses its published Spark 1.2 Standard example rather than assuming the newest model's full pricing. Qwen3.8 cache rates require console confirmation. BytePlus shows Seed 2.1 Turbo input/output pricing publicly. Gemini cache storage is additional. DeepSeek's off-peak price is 50% of peak; all other hours outside its weekday peak windows qualify. [5–11]

API-equivalent value is not provider cost. Retail price contains margin; a subscription may use different scheduling, batching, tool infrastructure and service priorities. More tokens also do not necessarily mean more useful completed work.

03 / Can the whole weekly allowance be converted?

Z.ai: a genuinely calculable weekly allowance

At the default 15% fresh / 80% cached / 5% output mix, GLM-5.3 uses 359.5 credits per million tokens. Max has 140,000 weekly credits. The half-price off-peak meter produces:

All peak
389.43M
$248.46 API value/week
All off-peak
778.86M
$496.91 API value/week
Monthly API value
$1,077–2,153
Not the subscription fee
Included constraint
28,000
Credits per five-hour period

Tools excluded; 5-hour constraints still apply. Max's actual checkout price was not verified. This is not a $200-plan claim. The official 95%-cache examples use a different workload and therefore show different token counts. [12]

04 / A 168-hour inference-demand model

Synthetic · not telemetry

Weekly average = 100 for EACH company. A value of 120 means 20% above that company's modeled weekly average, not 120% GPU utilization.

Move over or tap the plot to inspect an hour. No live data is collected.

A flat baseline plus regional work/consumer cycles. Geographic shares and amplitudes are assumptions; no company hourly request data was available. Training may fill inference troughs, so this is NOT a chart of total compute electricity or fleet occupancy.

Provider / groupBusy window in Singapore timeEvidence quality
OpenAI / AnthropicRoughly weekday evenings → 20:00–02:00Inference; historical Claude promotion is weak corroboration
Gemini / Grok / MetaBroader global peaks; less confident daily timingAssumed mixed work/consumer footprint
DeepSeekMon–Fri 09:00–12:00 and 14:00–18:00Published current peak-price windows [7]
Z.aiMon–Fri 14:00–18:00Published current peak-credit window [12]
Qwen / Kimi / ByteDanceAsian workday; consumer demand can extend into eveningAssumption, not a disclosed usage profile

US daylight-saving time is in effect for this September snapshot. Do not interpret the model's highest single hour as a measured “busiest hour.” A launch or outage can overwhelm these ordinary-week patterns. Historical Claude corroboration refers to March 2026, not a current price policy. [27]

Transparent assumptions for all ten curves
DeveloperUS East proxyEurope proxyAsia proxyWork-oriented fraction

Each region uses a daily Gaussian work profile centered at 13:00 (width 4h) and a consumer profile centered at 20:00 (width 4h). Regional weekday weights: 1, 1.03, 1.05, 1.03, 0.95, 0.68, 0.63. Consumer weekend weight: 1.03. Raw load = 0.35 + 0.65 × weighted regional cycles; then normalize the weekly mean to 100. All coefficients are illustrative.

05 / Monthly spending: two different accounting views

Illustrative genAI operating-resource cost, US$ billions per month. For conglomerates, exclude ordinary advertising, retail, legacy cloud and non-genAI businesses.

Do not add the infrastructure column to total cost. It is a cross-cutting subset already distributed across inference, research and operations. New hardware purchases are capex; rental and depreciation are operating-resource costs. Cash burn is neither of those totals.
InferenceResearch / trainingOther

Lengths show central scenarios. Every value is modeled, not disclosed.

Evidence that anchors the scale

OpenAI: reported $40B revenue run rate; reported 2025 adjusted gross margin of 33%. These do not reveal today's spending. [17, 18]

Anthropic: reported $65B run rate at July-end. Its earlier Q2 forecast implied roughly $3.45B monthly adjusted operating expense, including training. [19, 20]

Compute rent: September reporting puts Anthropic's SpaceX agreement at $1.25B/month. This is one supplier commitment, not the whole cost base. [21]

Z.ai: H1 R&D of CNY2.1B implies roughly $52M/month at the report's exchange rate; growth can raise the later run rate. [25]

For private labs and unreported genAI segments, remaining allocations are scenario judgments, not fitted audited results.

DeveloperTotal baseBroad rangeInferenceR&D / trainingOtherInfra subsetConfidence

Research includes pretraining, post-training, reinforcement learning, synthetic data, evaluation and research staff. Infrastructure includes rented capacity and owned-fleet depreciation, power and networking; the own-versus-rent split is not observable. Excludes new capital purchases, acquisitions, financing and stock-compensation revaluations. Ranges are plausible sensitivity bands, not statistical intervals.

Actual or guided infrastructure purchases — separate from the model above

OrganizationReported / guided capexMonthly equivalentScope warning
Alphabet$195–205B / full-year 2026$16.25–17.08BWhole company, not just Gemini [22]
Meta$130–145B / full-year 2026$10.83–12.08BWhole company, not just Muse [23]
SpaceX AI infrastructure$16B / Q2 2026$5.33BIncludes capacity sold to other labs [21]
AlibabaApproximately $10B / June quarterApproximately $3.33BBroader cloud / company capex [24]
Company-by-company reasoning and limitations

06 / Who is profitable to serve?

Estimated direct-delivery gross margins, not audited segment results. Research, sales and fixed corporate overhead are excluded from this metric.

Gross margin = (recognized customer revenue − directly attributable delivery cost) / revenue
DeveloperRetail API scenarioBusiness / enterprise scenarioTypical paid-user scenario

These are low-confidence economic priors, not measured differences between companies. Enterprise ranges assume a mix of seats, usage billing and support. Discounted bulk API may be less profitable. “Typical” deliberately excludes customers exhausting every allowance. Meta refers to commercial Standard/Muse Code offerings, not ad-funded free assistants. Unverified subscription cohorts are not estimated.

The maxed-out $200 subscriber

Take the illustrative $300/week workload above: $1,300/month of API list value. Let direct compute cost equal r% of that list value and other delivery cost be $10/month.

Cost fraction rTotal costGross profitMargin
10%$140+$60+30%
20%$270−$70−35%
35%$465−$265−132.5%

Break-even r ≈14.6%. This demonstrates the threshold; it does not establish OpenAI's actual cost fraction or your allowance.

Why the business can still work

Unused allowances: a $200 subscriber costing $10–40 to serve has an 80–95% direct margin.

Metered consumption: API bills normally rise with costly usage; a flat subscription fee does not.

Utilization and caching: stable workloads and reusable context can make the same physical fleet cheaper per useful task.

Product constraints: model sublimits, rolling windows, off-peak incentives and paid overages control expensive tail usage.

Portfolio economics: a loss-making heavy-user cohort can coexist with positive overall delivery margin. That still does not pay for all frontier-model research.

Illustrative OpenAI bridge: $40B annualized revenue ≈ $3.33B/month; less $2.2B modeled inference delivery gives about $1.13B delivery contribution. Less $3.0B research and $0.8B other costs gives about −$2.67B/month operating contribution. This combines a reported revenue snapshot with assumptions; it is NOT a reported income statement.

Provider “adjusted gross margin” definitions may allocate costs differently. Reuters also notes differing treatment of partner revenue shares, which limits comparisons. Zero-revenue free accounts have negative gross-profit dollars if costly to serve, but their percentage gross margin is undefined; ads or ecosystem benefits require separate allocation. [20]

Evidence & methodology

Primary documents for pricing and limits; reported company finances where audited segment disclosures are unavailable. Checked 13 September 2026. Dynamic pages may later change.

What this research does not establish

It does not measure a Pro account's actual weekly credit pool, a vendor's real internal GPU-hour cost, individual customer gross margin, each company's hourly traffic, or audited current genAI-only operating expense. Some dynamic checkout and cache-price pages were inaccessible; these fields remain unverified. The calculator makes assumptions explicit, and the demand curves are a synthetic model rather than fabricated observations.

Retail prices alone do not rank task quality, reliability, privacy, wall-clock completion time or total project cost. A cheaper token can still be more expensive per successful task. Subscription credits are not transferable API credit and cannot be redeemed for the indicated API-equivalent value.