A textbook · Manuscript, 2026

Production, Pricing, and Value in Token-Metered Intelligence

Quanyan Zhu · New York University

The token is the kilowatt-hour of this new utility.

Machine intelligence is becoming a metered service, like electricity or bandwidth. This book studies the economics of that service: how tokens are produced by inference hardware, how applications and agents consume them, what the spending buys, and how the resulting market should be priced, governed, and insured.

21chapters
7parts
4appendices
~710pages
Cover of AI Tokenomics by Quanyan Zhu
Why the economics comes next

Deployment has outrun understanding

For an enterprise

Deciding whether to roll out an agent means knowing what it will consume, how that consumption grows as tasks lengthen, what it will cost at the prices actually paid, and whether the outcomes are worth more than the tokens.

For a provider

Pricing a service whose cost depends on hidden reasoning, and whose value depends on results the provider never observes.

For a regulator

Telling apart the features of the market that concern competition, those that concern risk, and those that are simply the arithmetic of a new utility.

For research

Systems papers stop at cost per million tokens; economics papers abstract away the machine that sets the cost floor. The book brings the two together around a single object: the token.

The organizing idea

From production to consumption to value

The book follows one chain, from how silicon, energy, and capital become tokens, through how applications and agents spend them, to what the spending buys. Pricing, markets, governance, and risk are built around it.

01 Production

The supply of inference

The cost of producing tokens, the throughput–latency frontier and optimal utilization, capacity procurement, hosting, energy, water, and power.

Part II · Ch. 4–5
02 Consumption

The demand for tokens

Context, caching, and retrieval; reasoning, agents, and branching consumption; an algebra for estimating and forecasting what a workload will cost.

Part III · Ch. 6–8
03 Value

From tokens to outcomes

Total cost and quality floors, routing and cascades, token allocation and shadow prices in agentic workflows, and the valuation and attribution of agents.

Part IV · Ch. 9–12
Pricing and markets (Part V, Ch. 13–16): price lists and menus; token-based, value-based, and hybrid contracts; demand, rebound, and market structure; markets for compute and for agents.
Governance, risk, and society (Part VI, Ch. 17–19): internal markets, budgets, and FinOps; risk, liability, and insurance; the political economy of tokenized machine labor.
Structural results

Four results that outlast any price list

Prices, products, and market facts change monthly, so the book keeps them in dated boxes. What it develops instead are structural facts that hold whatever the current prices are.

01 · Production

Utilization enters unit cost hyperbolically

A pool of hardware costs the same per day whether or not it is busy. If it costs Cf a day, can deliver Nmax tokens, and runs at utilization u, the capacity cost of a delivered token is inversely proportional to u. On a dedicated pool, cutting tokens by 90% raises the cost per token tenfold.

c̄(u) = Cf / (u · Nmax)
02 · Consumption

Agentic consumption can diverge

When an agent delegates, each invocation can spawn sub-investigations, and consumption becomes a branching process with mean branching factor m. At m = 0.8 an alert costs 5 calls on average; at 0.95, 20 calls; at m ≥ 1 the mean is no longer finite.

E[calls] = 1 / (1 − m), for m < 1
03 · Consumption

Replayed context grows quadratically

Every call of an agent loop re-sends the history. In the book's example, ten rounds consume 65,000 input tokens and forty rounds consume 860,000: four times the rounds, 13.2 times the tokens. Try it below.

input tokens grow like n²
04 · Value

Token, cost, and value savings differ

Token savings are not automatically cost savings: on a dedicated pool the daily bill is unchanged until capacity shrinks, and cheaper tokens invite more use. Cost savings are not value either, since value is not monotone in cost or in any single quality metric.

Δ tokens ≠ Δ cost ≠ Δ value
Interactive · Result 03

Why long agent sessions get expensive

An agent that keeps its conversation in context re-sends everything it has seen on every turn. Change the settings and watch the cost curve bend upward. Prompt caching lowers the price of the replayed part, but the quadratic shape remains.

Workload
Prices (hypothetical)

Prices are illustrative, not any provider's list price. The model: each turn re-sends the system prompt and the full history, then adds new input and generates output.

–session cost, no caching
–session cost, with caching
–last turn vs. first turn
–of input is replayed history
Cumulative session cost by turn
No caching With prompt caching If every turn cost the same as the first
Tin(n) = n·S + n·u + (u + o)·n(n − 1)/2
Who it is for

Six readers, six paths through the book

Those who price AI services

The cost structure behind a price list, menus and screening, and token-based versus outcome-based contracts.

Parts II and V

Engineers and architects

Serving cost, consumption, caching, routing, and workflow allocation, turned into dollars and quality.

Parts II–IV

Students and researchers

Worked examples, exercises, offline laboratory exercises, and a map of open problems across fields.

Ch. 21 and Appendix C

Policy makers and regulators

Market structure, competition, liability, insurance, compute governance, and labor, separated from transient prices.

Ch. 1, 3, 9, 15, 16, 18, 19, 21

Finance, procurement, FinOps

Total cost of ownership, the choice of cost denominator, chargeback and budgets, and an enterprise program from visibility to value.

Ch. 9, 17, 20

Managers and executives

The questions to ask before a rollout, metrics that reward the right behavior, and the risks a token bill does not show.

Ch. 1, 9, 13, 17, 20
Running examples

Three systems carried through every chapter

The second example is calibrated on a real experiment of 480 metered runs, so its numbers are measured rather than assumed. Every worked number in the book is reproduced by an accompanying script.

R1

Customer-support assistant

A high-volume, low-margin service where utilization, caching, and routing decide the unit economics.

R2

Invoice-reconciliation agent

A multi-step agent whose cost depends on reasoning depth and retries, measured across 480 metered runs.

R3

Multi-agent security operations

A team of agents triaging alerts, where allocation, attribution, and liability all come into play.

Contents of the complete manuscript

21 chapters in 7 parts

Part I is available now as a free sample. Every chapter opens with learning objectives and closes with key takeaways, notes on the literature, and exercises.

IFoundations: Tokens as an Economic Resourcein sample
  1. 1The Economic Turn in Machine Intelligence
  2. 2How Inference Works: A Primer for Economists and Managers
  3. 3The Token as a Unit of Account
IIProduction: The Supply of Inference
  1. 4The Cost of Producing Tokens
  2. 5Capacity, Hosting, and Energy
IIIConsumption: The Demand for Tokens
  1. 6Token Demand: Context, Caching, and Retrieval
  2. 7Reasoning, Agents, and Branching Consumption
  3. 8Estimating Workload Cost: From Big-T to an Expected-Cost Algebra
IVValue: From Tokens to Outcomes
  1. 9Total Cost, Outcomes, and Value
  2. 10Routing, Cascades, and the Deployment Problem
  3. 11Token Allocation in Agentic Workflows
  4. 12Valuing and Attributing Agents
VPricing and Markets
  1. 13Pricing Token-Metered Services
  2. 14Token-Based, Value-Based, and Hybrid Contracts
  3. 15Demand, Rebound, and Market Structure
  4. 16Token Markets and Agent Economies
VIGovernance, Risk, and Society
  1. 17Internal Markets, Budgets, and FinOps
  2. 18Risk, Liability, and Insurance
  3. 19Agentic Capital and the Political Economy of Tokens
VIIPractice and Prospect
  1. 20Case Studies in AI Tokenomics
  2. 21Open Problems and a Research Agenda
App.Appendices
  1. AProbability, Queueing, and Optimization
  2. BEconomics Primer
  3. CLaboratory Exercises (run offline, no provider account needed)
  4. DGlossary
Related research

Papers behind the book

The book is the teaching treatment of a framework developed in the following research.

  • AI Tokenomics: The Economics of Tokens, Computation, and Pricing in Foundation Models
    Q. Zhu · arXiv:2606.24616, 2026
  • Agentomics: Economic Foundations for the Valuation, Attribution, and Pricing of AI Agents in Human-AI Workflows
    Q. Zhu · arXiv:2606.14769, 2026
  • PACT: A Contract-Theoretic Framework for Pricing Agentic AI Services Powered by Large Language Models
    Y.-T. Yang and Q. Zhu · IEEE GLOBECOM 2025 · DOI
  • Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs
    Y.-T. Yang and Q. Zhu · arXiv:2605.23929, 2026
  • Insurance of Agentic AI
    Q. Zhu · arXiv:2606.05449, 2026
  • AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation
    Q. Zhu · arXiv:2607.13230, 2026

A note on the word. “Tokenomics” was first used in cryptocurrency, for the supply schedules and incentives of digital assets. That usage is unrelated to this book beyond the shared word: here a token is the unit in which a language model reads and writes, and the unit in which its work is billed.