EN / 中文 ↗
v0.6 · §4 added (CapEx vs Revenue + Neocloud revenue structure: renting GPUs vs selling tokens)

Token Economics in the AI Era

From physical compute to the end user: a 4-layer value chain stratified by "Unit of Trade"

Research: Jimmy · 2026-05-10 · All chips linked to wiki entities (effective after deployment)

🆕 Sister panel: AI Application-Layer Funding · 63 companies / 13 domains (incl. full inference & neocloud funding history) ↗

The through-line of this research is the production and consumption of tokens. The ultimate goal: from underlying compute to end-user price, see clearly how value is distributed at every link in the chain and who takes the lion's share.

The key change in v0.2 was establishing an organizing principle: each layer uses its "output unit of trade" as the stratification marker. This turns "layer" from a "role label" into a "rigid criterion" — same transaction unit means same layer.

§ 1Organizing Principle: Unit of Trade Defines the Layer

What defines each layer is not "what it does" but "what unit of trade it outputs." The unit of trade converts between layers; the conversion point is the layer boundary.

Layer Input Key Conversion Output (Unit of Trade)
L1 Physical Compute Infrastructure + Chips Deployment + Integration GPU + Rack (physical hardware)
L2 Infra GPU + Rack Deployment + Operations GPU Hours (machine-hours)
L3 Model & API GPU Hours Training (→ Model Weights) + Inference Token API (per M tokens)
L4 Consumption Token API Wholesale / Direct Sale / Application Integration Token API (Path 1·2 → developers)
Subscription + Credit (Path 3 → end users)

§ 2Value Chain Framework

Top to bottom, arrows label the unit of trade passed between each layer. L3 and L4 each have an internal sub-structure (1P/3P, wholesale/retail) — expanded below.

L1 Physical Compute Physical Infrastructure

The bottom layer of hardware reality — infrastructure (data center / power / networking) + the AI chips themselves. Outputs "physical carriers of compute."

AInfrastructure

Physical space, stable power, high-speed interconnect, and electrical & cooling equipment — the AI factory's "land + nervous system."

A.cPower Distribution & Cooling
▸ Electrical (UPS / PDU / switchgear)
▸ Liquid cooling · DLC and cold plates (independent players almost entirely acquired by industrial conglomerates)
▸ Liquid cooling · Immersion (independent specialist players)
▸ More long-tail players (Legrand · Generac · Kohler · JetCool · Chilldyne, etc.) + full M&A landscape and bottleneck cascade analysis see L1 Physical Compute Deep-Dive Module
Baseline · A 1 MW AI DC total construction cost: ~$20–30M (high-density, AI-optimized, including full-stack electrical + cooling + networking equipment) · PUE 1.1–1.3 · Electricity price $0.05–0.12 / kWh · Depreciation 20–30 years · AI rack density 30–100+ kW (traditional ~10 kW)
Sub-itemShareKey Suppliers
Electrical (UPS / PDU / switchgear / distribution)40–50%Vertiv · Schneider · Eaton · ABB
Cooling (mostly liquid)15–20%Vertiv · CoolIT · nVent · Boyd
Networking (switches + optical modules + NIC)12–18%Arista · Broadcom · NVIDIA · Coherent · Innolight
Building shell15–20%Colocation providers + general contractors
Design / management / permits5–10%
Sources: JLL / CBRE / Synergy Research / Wood Mackenzie / Dell'Oro · Vertiv·Schneider·Arista·Broadcom·Coherent·Innolight Q1 2026 earnings (composite estimate)
BChips (AI Accelerators)

Turn silicon into AI compute units, sold or self-used in the form of "cards" or "systems."

B.bGoogle TPU (in-house, self-use + external via GCP)
B.cCloud-provider custom chips (Amazon · Microsoft, in-house + partial external)
Baseline · B Representative chip single-card pricing (2026 estimate): NVIDIA H100 ~$25–40k · B200 ~$30–50k · GB200 NVL72 rack ~$3M (72 cards) · AMD MI300X ~$20–30k
Per-card power: H100 700W · B200 1000W · GB200 superchip ~2700W · HBM: H100 80GB · H200 141GB · B200 192GB
Depreciation period: 5–6 years (per hyperscaler 10-K filings) — vs infrastructure 20–30 years. Over a single DC's lifetime, you replace GPUs 4–5 times. Sources: MSFT / AMZN 10-K (depreciation) · SemiAnalysis · secondary-market reporting
Upstream ▲ The physical manufacturing of every B-class chip depends on wafer foundries — the "L0" layer beneath L1. TSMC · Samsung Foundry
Synthesis · 1 MW Capacity Combining A infrastructure + B chips + power: full-stack annualized cost of 1 MW AI capacity.
ItemCapital InvestmentDepreciation PeriodAnnualized Cost
Infrastructure (DC building)~$15M25 yr~$0.6M
GPUs (~750 H100 @ $35k)~$26M5 yr~$5.2M
Power (1 MW × $0.08/kWh × 8760h)~$0.7M
Total~$41M~$6.5M / year
Inference: physical cost floor of a single H100·hour ≈ $1.0 / hr ($6.5M ÷ 750 cards ÷ 8760h). Within that, GPU depreciation accounts for ~80%, infrastructure + power together ~20% → L1 total cost is dominated by the chip.
vs L2 New Cloud (CoreWeave / Lambda) H100 reserved ~$2–3/hr → markup 2–3×, gross margin ~50% (consistent with public earnings disclosures).
Unit of Trade
GPU + Rack
Sale / Lease · one-time capital
L2 Infra Layer Providers / Operators

Directly owns and operates L1 assets, converting hardware into "machine-hours" sold externally.

Unit of Trade
GPU Hours
$/GPU·hour · billed by machine-hour
L3 Model & API Layer Model & API

Converts GPU Hours into a callable Token API. Internally formed by the convergence of two streams: training and inference.

Internal conversion flow
Input: GPU Hours
Training
Burns corpora into model weights. A one-off sunk cost, often tens to hundreds of millions of dollars.
→ Model weights (open / closed source)
Inference
Turns weights × GPU into real-time token capacity. Every call is a marginal cost.
→ Inference capacity
▼ Merge ▼
Model weights + inference capacity = Token API
Output: Token API (per M tokens)
Player taxonomy: 1P vs 3P have radically different cost structures

Both kinds ultimately "output a Token API," but the costs they bear are entirely different — a critical cut that any subsequent "value-chain verification" must make cleanly.

L3a · 1P (self-trained + self-operated)

Bears training sunk cost + in-house inference marginal cost. Strong pricing power, must amortize training investment.

L3b · 3P (deploying others' weights)

Bears only inference cost (open weights are free; licensed weights are paid). Essentially Inference-as-a-Service.

Engines powering 3P: Inferact (vLLM, $800M)RadixArk (SGLang, $400M)NVIDIA TensorRT-LLM

Baseline · DeepSeek Use DeepSeek as the L3 break-even calibration point. Open weights + extreme compression (MoE architecture + in-house inference optimization) let it, while consuming an equivalent amount of underlying GPU resources, push API pricing down to ~$0.14 / $0.28 per M token (input / output).
Usage: treat DeepSeek as the break-even floor under "zero brand premium + zero training amortization + inference-efficiency ceiling." The spread between any other 1P / 3P pricing and DeepSeek = the sum of that provider's brand premium + training amortization + inference-efficiency loss. Subsequent L3 markup calculations use DeepSeek as the calibration anchor.
Unit of Trade
Token API
$/M tokens · in / out billed separately
L4 Consumption Layer Token Consumption

The Token API is consumed at this layer. After entering from L3, it splits into three parallel paths flowing to different endpoints — some continues to circulate as Token API among developers; some is packaged by application-layer companies into "Subscription + Credit" sold to end users.

Input: Token API (from L3)
Path 1 Wholesale

Via an Aggregator or Token Brokerage, N 1P / 3P APIs are aggregated behind a single interface and then resold at a markup.

Output Token API (with markup) Other developers
Path 2 Direct Sale

An AI Lab (1P) or a 3P provider sells the Token API directly to developers, with no middleman markup.

Output Token API (direct) Developers
Path 3 Application Integration

The Token API enters application-layer companies, where it is wrapped into vertical products by industry. These companies ultimately sell to end users as "Subscription + Credit" — the highest-markup segment of the entire value chain happens here.

General Chatbot ChatGPT Claude.ai Gemini Poe
Coding Cursor Windsurf GitHub Copilot Cognition · Devin Replit
Legal Harvey Robin AI Spellbook
Presentation / Slides Gamma Tome Beautiful.ai Prezi Minds
Customer Service Crescendo Decagon Sierra
Search Tavily Perplexity You.com
Sales / Marketing Attentive Postscript Jasper Copy.ai
Finance Surf
AI + Payment Coinbase AgentKit Stripe · Tempo MPP Skyfire Kite AI x402 Protocol
Multi-channel IM Tinyfish Locus
Personal Agent Magic Character.ai Pi
Notes / Writing Notion AI Granola
Output Subscription + Credit End user

§ 3Vertical Integration

The unit of trade defines crisp layer boundaries, but the leading companies frequently swallow up adjacent layers, folding upstream and downstream into a single house. When doing "value-chain verification," you must first separate these companies' financials from any single-layer markup calculation — otherwise gross margins get blended into a single pot.

Direction Representatives Economic Implication
L3 → L1
AI Lab builds own compute
OpenAI Stargate · Google TPU+DC · Meta 350k H100 · xAI Colossus Big labs skip L2 and procure bare metal / build their own data centers directly. L2 New Cloud's bargaining ceiling is pushed down.
L2 → L3b
Cloud providers move up to API
AWS Bedrock · Azure OpenAI · GCP Vertex L2 uses its distribution to package closed-source models as 3P APIs and resell them, blurring the 1P / 3P boundary and absorbing part of L3 profit.
L2 + L3b integrated
Independent 3P with in-house compute
Together AI · Fireworks · Groq A single company straddles L2 / L3b. Token cost cannot be cleanly separated in the books, and external pricing does not expose per-layer markups.
L1 → L3 / L4
NVIDIA software tax
NVIDIA NIM · DGX Cloud A hardware vendor jumps over L2 and goes straight to inference services and API platforms. CUDA lock-in extends upward.
L3a → L4 Path 3
1P self-operated end-product
OpenAI ChatGPT · Anthropic Claude.ai · Google Gemini app An AI Lab runs its own end-product, bundling 1P API + application under one roof. Gross margin is not observable from outside.

§ 4Demand vs. Supply: CapEx vs Revenue (2026 Q1)

The first three sections settled "how the layers divide and who swallowed whom." This section answers a plainer question: how many times larger is the money the whole industry pours in than the money it actually earns back? Data is based on 2026 Q1 earnings (disclosed 2026-04-29) and public commitments; all amounts in USD.

2026E Global AI economy (annualized) Amount (B = billion USD) Note
Global AI infra CapEx, total ~$1,050B Big 5 + China + Neocloud + Stargate + Sovereign AI
Big 5 hyperscaler 2026 CapEx guidance (combined) ~$725–760B AMZN $200B · MSFT $190B · GOOGL $185B · META $135B · ORCL $50B
Gross AI end revenue (Labs + Cloud AI + Neocloud) ~$150–160B before stripping circular flows
Circular-flow estimate (Labs → Cloud return) −$28B (range $21–35B) OpenAI→Azure · Anthropic→AWS+GCP · others
Net external AI revenue (after de-circulation) ~$110–120B the real "new money" entering the ecosystem
CapEx / Net-revenue ratio ~9–10× vs fiber bubble ~2× · cloud buildout ~2.7×

Sources: Q1 2026 earnings calls (CNBC / Fortune / Microsoft Source / Alphabet earnings, Apr 29 2026) · Synergy Research (Apr 2026) · Sacra · CreditSights · Morgan Stanley · The Information.

A. Supply side: Big 5 CapEx, 5× in 4 years

Hyperscaler CapEx traced a near-straight line upward across 2023–2026. From the ~$155B range in 2023 to a 2026 guidance median of $735B, with ~60–70% annualized growth sustained. All five had Q1 CapEx YoY growth above 65%; Oracle tripled in a single quarter.

trendBig 5 CapEx, annual trajectory
FYBig 5 combined CapExYoYContext
2023$155BFirst year after ChatGPT; still normal levels
2024$256B+65%First acceleration wave
2025$443B+73%Training compute rolled out at scale
2026E$725–760B+64–72%Current guidance; several raised within the year
2027E~$1,050B+40%Goldman / Moody's forecast; may cross $1T within the year

CapEx / Sales ratio at record highs: Oracle 86%, Meta 54%, Microsoft 47%, Alphabet 46%, Amazon 25%. Big 5 free-cash-flow coverage dropped below 1.0× for the first time — CapEx + dividends + buybacks now exceed operating cash flow.

Source: Q1 2026 earnings + Introl / Morgan Stanley estimates.

B. Demand side: ARR structure split by layer

Reorganizing revenue by this token map's 4-layer framework shows "which layer the money enters from." Cloud AI revenue (L2+L3b) is currently the largest ($74.5B), but much of it comes from AI Labs' compute return flow; AI Labs (L3) external revenue ~$57B, Neocloud (L2) ~$10B.

Layer Representatives + scope 2026E ARR Key watch
L3 · AI Labs Anthropic + OpenAI + xAI + Mistral + Cohere + Perplexity + others ~$57B Anthropic $30B (disclosed Apr, only $9B at year start) overtakes OpenAI (~$25B) for the first time; Top 2 = 96%
L2/L3b · Cloud AI MSFT AI Business + AWS Trainium/Bedrock + Google Cloud GenAI ~$74.5B MSFT $37B (incl. OpenAI usage, +123% YoY) · AWS $20B · GCP ~$17.5B (+800% YoY)
L2 · Neocloud CoreWeave + Lambda + Together + Crusoe + Nebius + others ~$10–25B CoreWeave $66.8B backlog · Nebius 6× growth; essentially swing capacity backstopping the two layers above
Total · gross revenue ~$150–160B net external ~$110–120B after stripping circular flows

C. Circular flows: three "investment ↔ compute consumption" ballast pairs

AI Labs' largest cost is compute, and the compute sellers are their largest shareholders. Investment buys compute commitments, compute commitments prop up valuations, valuations buy more investment — this loop means cloud AI revenue's "gold content" needs a discount. Three main loops:

loop 1Microsoft ⇌ OpenAI
MSFT cumulative investment in OpenAI$13B+ (incl. 27% profit-share equity, ~$135B mark)
OpenAI Azure long-term commitment$250B (~7 years, ~40% of MSFT's $627B RPO)
OpenAI 2026E ARR~$25B (year-end expectation $30B+)
Est. annualized Azure return~$10–20B/yr (50–80% of OpenAI revenue)

MSFT's "AI Business" $37B ARR disclosure scope itself states it "includes model builders' Azure usage" — i.e., OpenAI's training + inference consumption is booked directly into Microsoft's AI revenue.

Source: Microsoft Source (Apr 29 2026) · om.co 10-Q analysis (May 1 2026).
loop 2Amazon ⇌ Anthropic
AMZN cumulative investment in Anthropic$8B (invested) + up to $25B (incl. convertibles), latest valuation $60.6B
Anthropic AWS long-term commitment$100B+ (10 years, 5GW Trainium compute)
Anthropic 2026E ARR$30B (disclosed Apr; $9B at year start → $30B Apr = 3.3× / 4 months)
Est. annualized AWS return~$10–25B/yr (Morgan Stanley est. 75% of cost is cloud)

AWS self-disclosed Trainium business $20B ARR, highly synced with Anthropic's commitment curve. A significant share of AWS's $364B backlog comes from Anthropic long-term contracts.

Source: Anthropic announcement (Feb 2026) · VentureBeat (Apr 2026) · Bloomberg.
loop 3Google ⇌ Anthropic
Google investment in Anthropic$10B invested + up to $40B milestones (incl. cloud credits)
Anthropic Google Cloud commitment$200B / 5 years (5GW TPU compute, starting 2027)
Share of GCP $462B backlog~40% (per $200B/$462B)
Est. annualized GCP return (mature)~$40B/yr (starts 2027; 2026 ~$2–5B)

GCP backlog doubled within Q1 2026 to $462B, driven precisely by Anthropic's $200B commitment (disclosed 2026-05-05). Google Cloud GenAI revenue +800% YoY, operating margin doubling from 17.8% → 32.9%, behind the same lock-in.

Source: The Information (May 5 2026) · Google Q1 2026 earnings call · BigGo Finance.

Circular-flow total: Anthropic + OpenAI alone have locked ~$550B in long-term compute commitments to the Big 3 clouds (AWS $100B + GCP $200B + Azure $250B), nearly half of the Big 3's (incl. Oracle) combined ~$2 trillion RPO. Annualized "double-counted" revenue est. ~$28B (range $21–35B), ~18% of gross AI revenue.

D. Historical comparison: how is this cycle different from fiber / cloud buildout?

Cycle Peak-year CapEx CapEx / Revenue Asset utilization Outcome
Fiber bubble 1998–2000 $120B/yr ~2× ~2% (measured 2002) WorldCom / Global Crossing etc. ~80% of companies bankrupt
Cloud buildout 2010–2015 $80B/yr ~2.7× rising Demand caught up; AWS/Azure/GCP became leaders
AI buildout 2024–2026 $1,050B/yr (2026E) ~6.6× gross / ~9.4× net ~97.5% (current GPU) Undetermined

Three data differences decide whether "this time is different" holds: (1) Speed — AI buildout is 3× faster than cloud buildout; (2) Concentration — 5 companies control ~60% of spend (vs the fiber era spread across dozens of telcos); (3) Utilization — GPU 97.5% vs fiber's 2% post-crash, the strongest "demand is real" evidence today, but also means no buffer.

E. Trend calls (5)

  1. 2027 CapEx crossing $1T is a "no-suspense" event.

    The Big 5 have already issued $108B in debt (3.4× the historical average); Morgan Stanley/JPM estimate $1.5T cumulative issuance across 2025–2027 to fund this wave. CapEx/Sales sits at historical extremes; future increments come from debt, not cash flow.

  2. The gross / net revenue gap will widen as circular commitments mature, not converge.

    Anthropic's $200B GCP contract doesn't book until 2027; OpenAI's 7-year $250B is still in its first half. Labs' return-flow share of cloud AI revenue is projected to rise from 18% to 25–30%, meaning "real revenue" growth lags the disclosed figures.

  3. Anthropic compressed "training cost / revenue" to 1/4 of OpenAI's — a new cost-curve watershed.

    Anthropic's 80× annual growth with training cost only 1/4 of OpenAI's means the efficiency frontier is shifting left. If frontier-model training cost structure enters a "train once, beats four prior" phase, 2027 CapEx ROI gets repriced.

  4. The consumer subscription / API inflection has arrived; enterprise penetration is accelerating.

    OpenAI 900M WAU (was 800M in 2025-10); M365 Copilot paid seats 15M → 20M in 6 months; Anthropic Fortune 100 penetration 70%, $1M+ annual-spend customer count doubled in 2 months to 1000+; Menlo Ventures measured enterprise LLM spend at 2.4× in half a year. Demand is shifting from "pilot" to "line-item IT budget."

  5. Neocloud is an observation point, not a sector.

    $10–25B ARR vs $40–65B CapEx is an inverse spread. CoreWeave's $66.8B backlog has Microsoft as its main customer (once ~62% of revenue) — essentially an off-balance-sheet extension of hyperscalers. Neocloud utilization / GPU spot price / hyperscaler in-house capacity are the early signals for whether this round is overheating.

F. Neocloud revenue structure: renting GPUs vs selling tokens

E5 treated Neocloud as a single ARR blob, but the same $10–25B mixes two very different kinds of money: renting bare GPU compute (compute rental, per GPU-hour) and selling tokens (inference API, per million tokens). This line decides whether a company looks more like a "commoditized GPU landlord" or a "sticky, high-margin software" business.

Company Rent GPUs · compute Sell tokens · API Scope / source
CoreWeave ~95%+ minimal Multi-year reserved + enterprise contracts; no standalone API; Microsoft once ~62% of revenue
Crusoe / Lambda ~90%+ small Energy-tied / developer-driven bare compute rental; H100 spot ~$3.9/GPU-hr
Nebius compute-dominated Token Factory (not broken out) Q1 2026 total revenue $399M (+684% YoY); 6-K reports a single "AI cloud" segment, inference revenue not disaggregated
Together AI ~60–70% ~30–40% ~$1B ARR; Sacra / Contrary estimate; the most token-leaning neocloud
SEC evidenceNebius 6-K: inference revenue is "not found" in the filings
Q1 2026 total revenue$399.0M (vs $50.9M, +684% YoY)
Segment scopeNebius AI cloud / Avride / TripleTen; the latter two "contributed only limitedly to group revenue"
Within AI cloudSingle revenue line (GPU compute + storage + software services), inference vs compute not split
Token Factory (launched 2025/11)Zero mentions in the entire 6-K

Implication: for neoclouds, "inference / Token Factory" remains narrative > financials — companies will tell the inference story on earnings calls, but in the revenue structure it is still too small to break out. Renting GPUs is still the base.

Source: Nebius Group 6-K Q1 2026 (SEC EDGAR, filed Apr 2026) · Sacra · Contrary Research.

Direction call: the closer to "renting GPUs," the more like a landlord (low margin, lives and dies with GPU spot price, bleeds the moment utilization dips); the closer to "selling tokens," the more like software (high margin, API stickiness). The industry trend is migrating toward the token end — Microsoft internal data: ~50% of GPU customers now access compute via AI APIs, only 20–25% still lock in reserved bare metal; inference is projected at ~2/3 of all 2026 AI compute (only 1/3 in 2023). Together is the only clearly token-leaning neocloud, the root of why its valuation narrative is more "software-like" than pure CoreWeave. Upstream engine-author companies (vLLM→Inferact $800M · SGLang→RadixArk $400M) entering managed-API hosting will pull value further from "renting GPUs" toward the "selling tokens" layer.

Scope & disclaimer: AI-specific CapEx is not separately disclosed by any company, estimated at the industry-standard 75%; circular flows are external estimates (companies do not disclose "OpenAI's spend on Azure"); some long-term commitments (Anthropic-GCP $200B) do not start booking until 2027. Core figures are as of 2026-05-11; §4·F (Neocloud revenue structure + Nebius 6-K) added 2026-06-28. Next review: 2026 Q2 earnings (~2026-07-29).

G. H100 rental price: three tiers, tracked weekly

§4·F argued that the more a neocloud "rents GPUs," the more it lives and dies with the GPU spot price — so here we pull the H100 $/GPU-hr out and track it weekly, split into spot / neocloud / hyperscaler (same chip, up to a 20×+ spread). The past-year storyline: a three-year slide bottomed in autumn 2025, then 2026 turned into a structural reversal — but it's the contract / tight tiers that reversed while marketplace spot stayed low. The divergence itself is the signal this panel watches.

Method & disclaimer: Once a week we take public on-demand single-H100 quotes from getDeploying, bucket providers into the three tiers, and take the median (self-built index, base 100 = first live week 2026-06-29). H100 composite avg = equal-weight mean of the three tier medians (dark-red line); its stat card shows price, index and week-over-week change (live points only, never across the seed/live seam). Dashed = reconstructed history (2025-07 to 2026-06, rebuilt from Silicon Data tier medians + spot-floor reports); solid = weekly live capture. The two baskets differ in methodology, so a level jump at the seam is expected — read each segment's trend, don't compare absolute levels across the seam. Dotted = published anchors of two paid industry indices (Silicon Data H100 blended index; SemiAnalysis 1yr-reserved contract index) — these are not scraped, just a few manually entered points disclosed in their blog/newsletter, overlaid to calibrate our self-built live line. Their methodology is not fully comparable to our three tiers.