No public, guaranteed weekly token allocation. Five-hour estimates are not a weekly cap.
Illustrative central case; $4–8B range. Not a disclosed company result.
Weekly GLM-5.3 tokens at a 15/80/5 mix; peak versus off-peak. [11, 12]
Scope: OpenAI, Anthropic, Google, xAI, Meta, DeepSeek, Alibaba, ByteDance, Moonshot and Z.ai. This is a selected group of major model developers, not a definitive valuation or intelligence ranking. All financial estimates are independent scenarios. USD prices exclude taxes, app-store differences and negotiated contracts.
01 / Your $200 Codex question
A reproducible conversion, with the unpublished variable left visible rather than invented.
Five-hour windows ≠ weekly tokens
Pro 20× documentation estimates 100–900 local Astra messages or 200–2,000 local Sol messages per five-hour window. Weekly limits also apply. Cloud chats use Sol. Message cost varies with work and context. [2]
Sol: 100 input · 10 cache · 500 output
At the default mix: Astra = 120 credits/M; Sol = 48 credits/M. If the actual included weekly balance is C credits, tokens in millions are C/120 and C/48. C is not published.
Output fixed at 5%; fresh input is the remainder. The $300 default is an illustration, not an estimate of your entitlement.
| Assumed weekly API value | Astra tokens | Sol tokens | Monthly API value | Value / $200 fee |
|---|---|---|---|---|
| $150 | 31.25M | 78.13M | $650 | 3.25× |
| $300 | 62.50M | 156.25M | $1,300 | 6.50× |
| $600 | 125.00M | 312.50M | $2,600 | 13.00× |
These three budgets are sensitivity points, not a probability range. Month = 52/12 weeks. “Tokens” includes repeated cached context and billable reasoning, not only visible output. Fast mode, long context, tools, cache creation and storage can change cost. If all 15% fresh input incurs 1.25× cache-write pricing, blended rates rise from $4.80/$1.92 to $5.175/$2.07 per million.
02 / What an identical token workload costs
Representative text models. Price per million tokens; standard short-context processing where applicable. Not a quality-equivalence ranking.
Swipe horizontally to inspect the full table.
| Developer / model | Fresh / M | Cache / M | Output / M | Cost / 100M | $200 API buys |
|---|
A dash means a rate was not independently verified here, not that caching is free or unavailable. Meta uses its published Spark 1.2 Standard example rather than assuming the newest model's full pricing. Qwen3.8 cache rates require console confirmation. BytePlus shows Seed 2.1 Turbo input/output pricing publicly. Gemini cache storage is additional. DeepSeek's off-peak price is 50% of peak; all other hours outside its weekday peak windows qualify. [5–11]
03 / Can the whole weekly allowance be converted?
Z.ai: a genuinely calculable weekly allowance
At the default 15% fresh / 80% cached / 5% output mix, GLM-5.3 uses 359.5 credits per million tokens. Max has 140,000 weekly credits. The half-price off-peak meter produces:
Tools excluded; 5-hour constraints still apply. Max's actual checkout price was not verified. This is not a $200-plan claim. The official 95%-cache examples use a different workload and therefore show different token counts. [12]
04 / A 168-hour inference-demand model
Synthetic · not telemetryWeekly average = 100 for EACH company. A value of 120 means 20% above that company's modeled weekly average, not 120% GPU utilization.
A flat baseline plus regional work/consumer cycles. Geographic shares and amplitudes are assumptions; no company hourly request data was available. Training may fill inference troughs, so this is NOT a chart of total compute electricity or fleet occupancy.
| Provider / group | Busy window in Singapore time | Evidence quality |
|---|---|---|
| OpenAI / Anthropic | Roughly weekday evenings → 20:00–02:00 | Inference; historical Claude promotion is weak corroboration |
| Gemini / Grok / Meta | Broader global peaks; less confident daily timing | Assumed mixed work/consumer footprint |
| DeepSeek | Mon–Fri 09:00–12:00 and 14:00–18:00 | Published current peak-price windows [7] |
| Z.ai | Mon–Fri 14:00–18:00 | Published current peak-credit window [12] |
| Qwen / Kimi / ByteDance | Asian workday; consumer demand can extend into evening | Assumption, not a disclosed usage profile |
US daylight-saving time is in effect for this September snapshot. Do not interpret the model's highest single hour as a measured “busiest hour.” A launch or outage can overwhelm these ordinary-week patterns. Historical Claude corroboration refers to March 2026, not a current price policy. [27]
Transparent assumptions for all ten curves
| Developer | US East proxy | Europe proxy | Asia proxy | Work-oriented fraction |
|---|
Each region uses a daily Gaussian work profile centered at 13:00 (width 4h) and a consumer profile centered at 20:00 (width 4h). Regional weekday weights: 1, 1.03, 1.05, 1.03, 0.95, 0.68, 0.63. Consumer weekend weight: 1.03. Raw load = 0.35 + 0.65 × weighted regional cycles; then normalize the weekly mean to 100. All coefficients are illustrative.
05 / Monthly spending: two different accounting views
Illustrative genAI operating-resource cost, US$ billions per month. For conglomerates, exclude ordinary advertising, retail, legacy cloud and non-genAI businesses.
Lengths show central scenarios. Every value is modeled, not disclosed.
Evidence that anchors the scale
OpenAI: reported $40B revenue run rate; reported 2025 adjusted gross margin of 33%. These do not reveal today's spending. [17, 18]
Anthropic: reported $65B run rate at July-end. Its earlier Q2 forecast implied roughly $3.45B monthly adjusted operating expense, including training. [19, 20]
Compute rent: September reporting puts Anthropic's SpaceX agreement at $1.25B/month. This is one supplier commitment, not the whole cost base. [21]
Z.ai: H1 R&D of CNY2.1B implies roughly $52M/month at the report's exchange rate; growth can raise the later run rate. [25]
For private labs and unreported genAI segments, remaining allocations are scenario judgments, not fitted audited results.
| Developer | Total base | Broad range | Inference | R&D / training | Other | Infra subset | Confidence |
|---|
Research includes pretraining, post-training, reinforcement learning, synthetic data, evaluation and research staff. Infrastructure includes rented capacity and owned-fleet depreciation, power and networking; the own-versus-rent split is not observable. Excludes new capital purchases, acquisitions, financing and stock-compensation revaluations. Ranges are plausible sensitivity bands, not statistical intervals.
Actual or guided infrastructure purchases — separate from the model above
| Organization | Reported / guided capex | Monthly equivalent | Scope warning |
|---|---|---|---|
| Alphabet | $195–205B / full-year 2026 | $16.25–17.08B | Whole company, not just Gemini [22] |
| Meta | $130–145B / full-year 2026 | $10.83–12.08B | Whole company, not just Muse [23] |
| SpaceX AI infrastructure | $16B / Q2 2026 | $5.33B | Includes capacity sold to other labs [21] |
| Alibaba | Approximately $10B / June quarter | Approximately $3.33B | Broader cloud / company capex [24] |
Company-by-company reasoning and limitations
06 / Who is profitable to serve?
Estimated direct-delivery gross margins, not audited segment results. Research, sales and fixed corporate overhead are excluded from this metric.
| Developer | Retail API scenario | Business / enterprise scenario | Typical paid-user scenario |
|---|
These are low-confidence economic priors, not measured differences between companies. Enterprise ranges assume a mix of seats, usage billing and support. Discounted bulk API may be less profitable. “Typical” deliberately excludes customers exhausting every allowance. Meta refers to commercial Standard/Muse Code offerings, not ad-funded free assistants. Unverified subscription cohorts are not estimated.
The maxed-out $200 subscriber
Take the illustrative $300/week workload above: $1,300/month of API list value. Let direct compute cost equal r% of that list value and other delivery cost be $10/month.
| Cost fraction r | Total cost | Gross profit | Margin |
|---|---|---|---|
| 10% | $140 | +$60 | +30% |
| 20% | $270 | −$70 | −35% |
| 35% | $465 | −$265 | −132.5% |
Break-even r ≈14.6%. This demonstrates the threshold; it does not establish OpenAI's actual cost fraction or your allowance.
Why the business can still work
Unused allowances: a $200 subscriber costing $10–40 to serve has an 80–95% direct margin.
Metered consumption: API bills normally rise with costly usage; a flat subscription fee does not.
Utilization and caching: stable workloads and reusable context can make the same physical fleet cheaper per useful task.
Product constraints: model sublimits, rolling windows, off-peak incentives and paid overages control expensive tail usage.
Portfolio economics: a loss-making heavy-user cohort can coexist with positive overall delivery margin. That still does not pay for all frontier-model research.
Provider “adjusted gross margin” definitions may allocate costs differently. Reuters also notes differing treatment of partner revenue shares, which limits comparisons. Zero-revenue free accounts have negative gross-profit dollars if costly to serve, but their percentage gross margin is undefined; ads or ecosystem benefits require separate allocation. [20]
Evidence & methodology
Primary documents for pricing and limits; reported company finances where audited segment disclosures are unavailable. Checked 13 September 2026. Dynamic pages may later change.
What this research does not establish
It does not measure a Pro account's actual weekly credit pool, a vendor's real internal GPU-hour cost, individual customer gross margin, each company's hourly traffic, or audited current genAI-only operating expense. Some dynamic checkout and cache-price pages were inaccessible; these fields remain unverified. The calculator makes assumptions explicit, and the demand curves are a synthetic model rather than fabricated observations.
Retail prices alone do not rank task quality, reliability, privacy, wall-clock completion time or total project cost. A cheaper token can still be more expensive per successful task. Subscription credits are not transferable API credit and cannot be redeemed for the indicated API-equivalent value.