Zero Times Infinity

Audience: Anyone trying to reason about where the AI money goes: engineers, operators, investors, and people who just opened their electric bill.
Reading time: ~15 minutes.

Two things I keep hearing about AI, often from the same person in the same conversation. The first is that it’s absurdly expensive: the datacenters, the gigawatts, the Nvidia market cap, the electric bill. The second is that it’s basically free now. A GPT-4-class answer that cost $30 per million tokens in 2023 costs a nickel today, and a frontier-class open model runs on a secondhand Mac in my office. Most of the confusion about AI economics comes from treating those as a contradiction. They aren’t. One is a price. The other is a bill. A bill is a price times a quantity, and the quantity is growing faster than the price is falling. The price is heading toward zero and the quantity toward infinity, and zero times infinity is not a number you can look up. Economists have a name for the pattern, Jevons’ paradox, from the observation in 1865 that more efficient steam engines made Britain burn more coal rather than less. What’s new is the number of zeros.

Some of the zeros: four companies will spend about three quarters of a trillion dollars on AI infrastructure this year. Nvidia is 8 percent of the S&P 500. The memory in your laptop roughly doubled in price. Google’s models process more than three quadrillion tokens a month. And in June the Bank for International Settlements put “AI capex bust” on its short list of things that could break the global financial system.

So this essay is about the bill: where it goes, who’s paying it, and how the two curves resolve. I’ll follow a dollar from a token back to a wafer, and on the way check the three intuitions I hear most often: that AI is driving up everyone’s electric bill, that the Nvidia-AMD-OpenAI-Oracle money-go-round has to end in a crash, and that there’s something called a RAM supercycle. Each is right in a way that’s more interesting than the slogan.

What a Gigawatt Buys

A year ago people counted AI in GPUs. Now they count it in gigawatts, and the switch tells you what became scarce.

A gigawatt is a billion watts, continuously: the average draw of about 800,000 American homes, and now the standard size of a serious AI campus. What one costs depends on who you ask. Jensen Huang, who has a certain interest in the number, said $50 to $60 billion in August 2025 and $80 to $100 billion this June for the next generation; Nvidia’s own accounting is that it collects about $25 billion per gigawatt of Blackwell and $40 billion per gigawatt of Vera Rubin, which is most of why the headline went up. Epoch AI built the most transparent model I’ve seen: a one-gigawatt Blackwell facility at $38 billion, of which $21 billion is servers, $11 billion is the building with its power and cooling, $5 billion is networking, and about $300 million is land and substations. Run it for a year and the fully loaded cost is around $8.5 billion, 60 percent of it the servers being written off.

Where the money goes in a one-gigawatt AI datacenter (Epoch AI model) One gigawatt, Blackwell generation (Epoch AI, May 2026) Build: $38B servers 56% building 30% net 13% Per year: $8.5B server depreciation 60% facility 16% net 14% electricity 7% Servers on a 5-year life, facility on 14. Land and substations are the 1% sliver at the far right of the top bar.

Keep that picture in mind, because it reorganizes almost every other argument. The building is real estate. The networking is plumbing. The chips are the money, and the chips are bought from one company at a 75 percent gross margin; Bernstein worked out that Nvidia’s gross profit alone is about 29 percent of everything anyone spends on an AI datacenter. Behind Nvidia sit TSMC, which fabs and packages every one of those chips, and the three memory makers who supply the high-bandwidth memory stacked on top. A gigawatt is a purchase order that flows to Santa Clara, then Hsinchu, then Icheon and Boise.

Now hold the unit up against the announcements. OpenAI’s chip deals alone come to about 26 gigawatts, plus 4.5 from Oracle, plus Sam Altman’s own framing of “about $1.4 trillion over the next eight years.” Anthropic’s commitments add up to something north of 13 gigawatts, with overlaps; Meta’s Hyperion site in Louisiana is designed for 5. At $40 to $60 billion a gigawatt, these are the trillions you keep reading about, and they are arithmetic on letters of intent. What’s energized is a different number. As of April, of Stargate’s seven US sites, only Abilene, Texas had power flowing, at most 0.6 gigawatts; the other six are foundations and steel due in late 2028, and the Abilene expansion was abandoned in March. Sightline Climate found nearly half the US capacity slated for 2026 delayed or cancelled, mostly over power, transformers, and neighbors. The first single-tenant sites reached roughly a gigawatt this summer, and Epoch measures the largest one doubling every ten months.

So the announced number is tens of gigawatts and the energized number is a few, and the gap is a power story. Chips are tight too; TSMC’s CEO said in July that packaging capacity “limits my customers’ growth.” But a chip queue is measured in quarters. A heavy-duty gas turbine ordered from GE Vernova today arrives in 2031.

The Power Bill

US datacenters used 4.7 percent of the country’s electricity in 2024, and Lawrence Berkeley National Laboratory’s central case for 2030 is 11.8 percent, a third of all US load growth this decade. Those numbers are real, and they’re where the intuition that AI is a power hog comes from. Per query, the intuition is off by orders of magnitude: Google measured the median Gemini prompt at 0.24 watt-hours, about nine seconds of television, and even a heavy reasoning query with 100,000 tokens of context runs to 40 watt-hours, less than two minutes of a hair dryer. The aggregate is right because the query count is in the quadrillions a month.

Here’s the part that changed how I think about all of this: electricity is not where the money goes. In Epoch’s model, the annual power bill is $594 million against $8.5 billion of total cost, 7 percent. If power were free, tokens would get about 7 percent cheaper. If the GPUs were free, they’d get 60 percent cheaper. That asymmetry explains behavior that otherwise looks insane: xAI running dozens of unpermitted gas turbines outside Memphis, Microsoft prepaying for substations. Nobody in this business is optimizing the power bill. They’re optimizing the number of days $20 billion of silicon sits in a crate. Satya Nadella said it plainly last November: “It’s not a supply issue of chips; it’s actually the fact that I don’t have warm shells to plug into.”

That’s also why “AI is raising my electric bill” is a fight about allocation rather than consumption. In PJM, the grid from Chicago to Washington, the capacity auction went from $29 per megawatt-day in 2024 to a regulatory cap around $330, where it has been pinned for three straight auctions, and the market monitor attributes $29 billion of the $64 billion those auctions collect to datacenter load. Nationally, though, the last five years of rate increases trace mostly to grid investment, wildfires, and gas prices, and Texas and Virginia, the two biggest datacenter states, had among the smallest. What decides the next five years is whether the people building gigawatts are made to pay for the turbines and transmission built to serve them. They can afford it; power is 7 percent of their cost. Whether they’re made to is politics, and the politics arrived this year: more than 300 datacenter bills in 30-plus states, large-load tariffs in 23 of them, fifteen gubernatorial candidates running on moratoria, and a Texas freeze on new grid connections in August after its interconnection queue hit 474 gigawatts. Wood Mackenzie reckons about 72 percent of US datacenter power requests are phantom.

And power is what binds the whole buildout. GE Vernova’s turbine backlog went from 100 to 116 gigawatts in a single quarter; the nuclear deals of 2024 and 2025 are mostly still paper, and if every one were built it would cover less than a fifth of projected 2035 datacenter demand. China added around 430 gigawatts of wind and solar in 2025 alone. The US added 53 gigawatts of everything.

The Memory Tax

The transfer from consumers to the buildout that actually happened in 2026 didn’t come through the meter. It came through memory.

The mechanism is a wafer trade. A gigabyte of high-bandwidth memory, the kind stacked next to every AI accelerator, takes three to four times the wafer area of ordinary DRAM, and AI will consume about a fifth of the world’s DRAM wafer output this year. The fabs chose the datacenter, and did it without adding wafer starts, because a new memory fab takes until 2028. Conventional DRAM contract prices rose a record 90 to 95 percent in the first quarter of 2026, then another 60 in the second; J.P. Morgan puts the cumulative rise since the start of 2024 above 400 percent. The DRAM industry’s quarterly revenue went from $27 billion in early 2025 to $155 billion this spring. SK Hynix reported a 76 percent operating margin for the June quarter, a point above Nvidia’s gross margin, and Samsung’s chip division expects to earn more this year than in its previous forty combined.

Everyone downstream paid. Apple raised Macs and iPads by $100 to $300, saying “we have never seen a component price increase this much, this quickly.” The console makers, Amazon’s devices, and Nvidia’s own desktop AI box followed; even Nvidia can’t get enough memory for Nvidia. Gartner has PC prices up 17 percent this year and shipments down 10. The Minneapolis Fed attributes about 0.4 points of core inflation to AI hardware demand, an effect it calls at least as large as tariffs. And the squeeze runs up the chain: memory was about 9 percent of the value of a Blackwell rack and is about a quarter of a Vera Rubin rack, Nvidia guided its gross margin down to 71 or 72 percent, and it told customers that systems shipping in early 2027 will cost 15 percent more. Three memory makers are now taking margin from the company that takes margin from everyone else.

In July the memory stocks crashed, 30 to 50 percent from their peaks, on fears that HBM expansion and a new Chinese entrant would swamp the market in 2027, in the same month Microsoft and Amazon raised capex guidance and blamed component inflation. Both were right; contract prices lag spot, and equities discount next year. Contract prices are still rising, at low-teens percent a quarter rather than 90, and the brokers who model it have supply and demand balancing in 2028.

I’ve written about the local-hardware side in Agentic AI @ 2 FPS: in this cycle, unlike 1998, the datacenter outbids the consumer for the wafer. If you want to know what AI cost ordinary people in 2026, don’t look at the electric bill. Look at the price of a laptop.

The Circle

On the last night of 1999 I was on Microsoft’s Redmond campus with a walkie-talkie. I ran a development team on MSN’s accounts and billing systems, and the radio net existed for the scenario where everything else failed at midnight. Nothing did. The other time zones had already rolled over, so by the time Pacific midnight arrived we were basically partying, and since you couldn’t tell who was talking, people were cracking jokes on a channel that had vice presidents on it. Possibly the vice presidents were partying too. There was no runbook to speak of; those were the wild-west years of online distributed systems, and next to what I’d see at Amazon sixteen years later it was comically naive, though par for the course. What I remember most clearly is that nobody on that channel said a word about the stock market. Ten weeks later the Nasdaq peaked, two weeks after that Cisco passed Microsoft as the most valuable company on earth, and by the end of March I’d left for a startup. It took me years to see the lesson: the people running the infrastructure never see the financial structure they’re inside. The circle gets drawn in a different room.

Here’s what that room looks like today. The deals first, because the shape is the argument.

In September 2025, Nvidia announced it would invest “up to $100 billion” in OpenAI as OpenAI deployed 10 gigawatts of Nvidia systems. It was a letter of intent; five months later no contract had been signed and no money had moved, and in March Huang said the $100 billion was “probably not in the cards.” What replaced it: $30 billion of straight equity in OpenAI’s $122 billion round at an $852 billion valuation, and then, in August, a guarantee capped at $105 billion backstopping OpenAI’s twenty-year leases on 4.25 gigawatts of datacenter at a SoftBank-owned campus in Ohio that, per Nvidia’s filing, “will exclusively host NVIDIA AI infrastructure.” Nvidia backstops the rent on a building that will only ever hold Nvidia chips, bought by a tenant Nvidia owns a piece of.

AMD signed OpenAI to 6 gigawatts of GPUs and issued it warrants for about 10 percent of the company, vesting as the gigawatts deploy, then signed the identical structure with Meta; two customers can now earn a fifth of AMD by buying AMD’s product. OpenAI’s largest single commitment is about $300 billion of cloud capacity from Oracle over five years, and Oracle is borrowing to build it: $43 billion of debt last fiscal year, another $40 billion of financing planned for this one, capex of $90 to $95 billion against operating cash flow of $32 billion, free cash flow of negative $23.7 billion. Most of that goes to Nvidia. OpenAI is about half of Oracle’s backlog; S&P cut Oracle to one notch above junk in July. Microsoft and Amazon close the loop from the other side, as the diagram shows, and Anthropic’s version has the same shape with Nvidia, Microsoft, Amazon, and Google. And in January Cerebras, a chip startup, took a billion-dollar working-capital loan to build capacity for OpenAI. The lender was OpenAI. A company that expects to burn $25 billion this year is financing its own supplier, which is 1999’s vendor financing run in reverse.

The circle: who funds OpenAI, what OpenAI owes them, and where the money goes next ...and buys Nvidia ...and buys Nvidia OpenAI valued at $852B, burns ~$25B/yr Nvidia $30B equity $105B lease guarantee orders 10 GW of systems AMD warrants: ~10% of AMD orders 6 GW Amazon $50B in, $138B AWS out Microsoft owns 27%, $250B Azure Oracle $300B contract, rated BBB- $300B over 5 yrs (Oracle borrows to build) green: money into OpenAI    blue: what OpenAI owes back    orange: where it goes next Anthropic's loop with Microsoft, Nvidia, Amazon and Google has the same shape.

Nvidia’s balance sheet is the clearest record of how far this has gone. Its equity investments went from $2.2 billion in 2024 to $99 billion in July, about $50 billion of it in the frontier labs. Guarantees total $108.5 billion. There is a $36 billion commitment to buy cloud capacity back from the AI clouds Nvidia sells chips to, which the filing calls a “new business model” and explains exists because those clouds “lack the ability to secure long-term infrastructure contracts and investment-grade financing capacity.” And in August Nvidia signed up Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR for financing platforms meant to mobilize $500 billion of other people’s capital, with Nvidia backing up to 25 percent of each deal. Huang: “In AI, compute is revenue.” Asked about the word circular, he called it “ridiculous.”

Is it circular? Partly. Nvidia’s cash into companies that buy Nvidia chips, set against about $300 billion of trailing revenue, is a single-digit to mid-teens percentage even if every dollar came straight back; Michael Burry’s “100 percent circular” is about forward commitments, not recognized revenue. The money flows one way, the way GM Financial’s does, so the deals don’t inflate today’s sales. What they do is correlate the risk. Every balance sheet in the loop now keys off the same variable, whether OpenAI and Anthropic keep raising money, and the contingent layer of guarantees and financing platforms is several times larger than the equity.

The precedent everyone reaches for is Lucent, which in 1999 had $8.1 billion of vendor financing outstanding on $38 billion of revenue. Forty-seven of those customers went bankrupt, Lucent’s revenue fell 69 percent in three years, and telecom bondholders recovered about twenty cents on the dollar. The difference in 2026, the one the bulls are right about, is that the customers have cash flow: Microsoft, Alphabet, Amazon, and Meta generated $451 billion of it in 2024, and the carriers of 1999 generated none. That difference is shrinking. The big four now spend roughly all of their operating cash flow on capex. Meta’s free cash flow fell 91 percent in the second quarter and it stopped buying back stock; Alphabet’s went negative and it raised $85 billion of equity in June, the first hyperscaler to sell shares for AI. Hyperscalers issued $182 billion of investment-grade bonds by mid-July, about 15 percent of all US corporate issuance. The “production capital, not financial capital” argument that made this feel different from 1999 was true in January. It’s fraying by the quarter.

Who Holds the Bag

So if the circle breaks, who’s holding what? Sort by tier, because the exposure is uneven and the credit market has already done the sorting.

Tier Who Exposure if the music stops
Hyperscaler equity Microsoft, Alphabet, Amazon, Meta Mild. Debt-to-equity in the single digits to low twenties, credit default swaps around 49 basis points. A bust is a capex cut and an earnings dip as depreciation catches up.
Intermediaries Oracle, CoreWeave and the neoclouds, private credit and SPVs Severe. Oracle at BBB- with negative free cash flow and half its backlog from one customer. CoreWeave with $35 billion of debt, paying $640 million of interest a quarter against $128 million of adjusted operating income, swaps near 855 basis points, roughly a coin flip on default within five years.
The labs OpenAI, Anthropic, xAI The bet itself. Equity-funded, about $190 billion raised between them in eighteen months at valuations near a trillion. OpenAI projects burning $25 billion this year and $57 billion next against a compute plan north of $650 billion through 2030.
Sovereigns and index funds Gulf funds, SoftBank, your 401(k) Structural. Nvidia is 8 percent of the S&P 500; the seven largest stocks peaked at 35 percent of it in June. Passive holders own this concentration by construction.

The middle tier is where 1999 lives. Meta’s Hyperion campus is financed through a special-purpose vehicle that issued $27 billion of A±rated bonds, a rating that rests on a Meta residual-value guarantee that doesn’t appear as debt on Meta’s balance sheet. Morgan Stanley estimates $1.5 trillion of the buildout through 2028 has to come from outside the hyperscalers’ cash flows, $800 billion of it private credit, and the BIS warned in June about deal terms “poorly disclosed, with risks of the same asset being pledged multiple times.” In July CoreWeave’s lenders repriced a $2.6 billion loan and added covenants; its stock fell 36 percent that month. Jim Chanos calls the neoclouds “financial conduits, not technology companies.”

The labs are the layer everyone else is lending against. When OpenAI’s CFO said in November that she’d like a federal “backstop” for AI infrastructure financing, and she and Altman walked it back within 48 hours, the interesting thing was how fast the market priced the implicit structure: Nvidia fell 7 percent that week. The closest thing to a real backstop so far is private: xAI’s 12.5 percent junk debt became investment grade in June by being folded into SpaceX and refinanced against Starlink’s cash flows.

Here’s the sort the credit market has made, in one table.

Borrower Five-year credit default swap, mid-2026
Amazon, Alphabet, Microsoft (composite) ~49 basis points, highest since 2018
Oracle ~200 basis points, an 18-year high, rated BBB-
CoreWeave ~855 basis points

The market thinks the risk sits with the intermediaries, not the hyperscalers or Nvidia. I think the market is right, with one caveat: the intermediaries are load-bearing. Oracle’s $300 billion, CoreWeave’s $22 billion, and the Ohio campus Nvidia is guaranteeing are a large share of the capacity OpenAI is counting on, and OpenAI is the customer Nvidia’s own filing describes as “a meaningful amount of our revenue” through those intermediaries. A default in tier two doesn’t stay in tier two.

The Price of the Frontier

The other half of the bill is what gets run on the capacity: training the models and serving them. Training first, because it’s the part with the most mythology.

The cost of the biggest training runs has grown about 2.4 times a year for a decade, by Epoch’s count. GPT-4’s final run in 2023 cost roughly $40 million in amortized hardware and power; the most expensive publicly estimated run to date is Grok 4, in mid-2025, at about $490 million and 310 gigawatt-hours. Dario Amodei’s ladder was $100 million, then $1 billion, then $10 billion by 2026, and as far as anyone can document no single run has crossed a billion; the $10 billion figures are program costs. Those are the real number anyway: the final run is about a tenth of a frontier lab’s research compute, and OpenAI’s research compute is projected at $19 billion this year. That’s also the right lens on DeepSeek’s famous $5.6 million: true, and the final pretraining run only, on a fleet SemiAnalysis costed at about $1.6 billion. Cheap final runs on expensive fleets are how the whole industry works; DeepSeek just said the number out loud.

Two shifts in 2026 change what “training” means. Post-training overtook pretraining: GPT-5 used less pretraining compute than GPT-4.5 while scoring higher, because the work moved to reinforcement learning. And the fast-follower discount got measured. Epoch puts open-weight models an average of four months behind the closed frontier, on an index where the frontier moves fourteen points a year, and the closed frontier trains on ten to twenty times the compute of the open one. Distillation is how: Anthropic reported one Chinese lab extracted more than thirteen million exchanges from Claude to train its own model.

Put those together and a frontier model is a depreciating asset with roughly a one-year half-life. A lab spends a billion-dollar program to hold a lead the commons matches in months at a tenth the cost. That’s why the labs race, why the training bill keeps rising even as last year’s intelligence goes free, and why the money has to be made during the months a model is ahead. Which is the serving business.

The Token Machine

The commodity tier did exactly what everyone predicted. GPT-4 launched in March 2023 at $30 per million input tokens; GPT-4o mini, which matches or beats it on most standard benchmarks, was $0.15 in mid-2024; GPT-5 nano is $0.05. That’s a 600-fold decline in thirty months at a fixed level of capability, the “tenfold a year” that a16z called LLMflation.

Price down, frontier price back up, volume up: 2023 to 2026, log scales Tokens: commodity price, frontier price, and volume (log scales) 2023 2024 2025 2026 $30 $3 $0.30 $0.03 per M input 10Q 3.2Q 1Q 100T 10T per month GPT-4, $30 GPT-5 nano, $0.05 cheapest GPT-4-class Anthropic flagship list price Opus 4.5, $5 Fable 5, $10 9.7T tokens/mo Google tokens per month Commodity price: cheapest OpenAI model that matches the original GPT-4 on standard benchmarks, input price. Frontier price: Anthropic's top model, input. Volume: Alphabet's disclosed monthly token counts.

But 2026 added a second curve, and it’s the one I’d want a board to understand: the price of the frontier went up. OpenAI’s flagship was $1.25 per million input tokens when GPT-5 launched in August 2025; the current GPT-6 Astra is $10. Anthropic went from $5 for Opus 4.5 to $10 for Fable 5. Google has announced its Flash prices double on January 1, 2027. Anthropic’s newer tokenizer produces about 30 percent more tokens for the same text, a price increase that never appears on a price page. So there are two curves now. Last year’s intelligence still gets ten times cheaper every year. This year’s got more expensive, because the labs discovered they have pricing power at the frontier, during the exact window the previous section says they have to monetize.

Volume is the other half of the identity, and it surprised everyone, including the people selling it. Google processed 9.7 trillion tokens a month in May 2024 and 3.2 quadrillion in May 2026, a 330-fold increase in two years. Epoch estimates token demand growing about ten times a year against inference capacity growing three to four times, which is a crunch by definition. Reasoning models generate around four times the tokens per query and now carry more than half of all routed tokens, and an agentic coding session is millions of tokens that run unattended.

Coding is where the volume comes from. Claude Code went from a $1 billion run-rate last November to roughly $8 billion by May; Anthropic’s total went from $9 billion at New Year to $65 billion at the end of July, and Amodei’s line was “we tried to plan very well for a world of 10x growth per year. And yet we saw 80x.” Coding leads because its return is legible: a verification loop turns tokens into merged, tested code, and a team can see the conversion rate. I’ve argued in The Bitter Lesson of Agentic Coding that the quality of that loop is the ceiling on what autonomous coding loops can produce; it’s also the reason the tokens get bought. Ramp’s card data shows the median company spending about $12 per employee per month on AI and the top 1 percent spending about $7,400. That 600-fold spread is the demand story in one number. Most companies are dabbling. A few have found a loop that converts tokens into money and are feeding it as fast as their vendors will let them.

The “sold below cost” worry, which I shared, turns out to be mostly wrong at the labs and right at the resellers. OpenAI’s margin on paid inference was reported around 70 percent and Anthropic’s API margin above 80; their blended gross margins are far lower because the free tier and the research sit in the same line. The tokens aren’t underpriced. Where inference really was sold below cost was one layer up: Cursor ran a negative 23 percent gross margin at the start of this year, repriced twice, and got acquired by SpaceX in June. And an H100 that rented for under $2 an hour at the end of 2025 rents for about $3.40 today.

So here’s the bull case in a single line: revenue is price times volume. For revenue to grow, volume has to outrun the price decline, and it has, by a wide margin. Every dollar of that revenue is somebody’s evidence that the capacity will be used.

How This Ends

Start with the historical shape. This is an installation period, in Carlota Perez’s sense: financial capital overbuilds the infrastructure, a crash transfers the assets to production capital at a discount, and the productive deployment happens afterward on infrastructure someone else paid for. Every analogy people reach for fits the pattern and was proportionally larger. Britain’s railway mania put 7 to 8 percent of GDP into rails in 1847 and a third of the authorized lines were never built. The telecom buildout put more than $500 billion into fiber, about 1.2 percent of GDP at the peak; 97 percent of it was dark in 2002, bondholders got twenty cents, and then the dark fiber became YouTube and Netflix and the cloud. AI capex is running at roughly 1.2 to 1.5 percent of US GDP. Jeff Bezos called this an “industrial bubble,” as opposed to a financial one, and meant it as a compliment: the industrial kind leaves things behind.

The twist is asset life, and it cuts both ways. Fiber lasts thirty years. A GPU is booked over five or six and is frontier-competitive for one to three; Burry’s claim that the hyperscalers are understating depreciation by $176 billion is a fight about whether the right number is three years or six, and the evidence is mixed. Either way a GPU glut clears in a couple of years; it isn’t the decade of free capacity dark fiber turned out to be. But look at what else is in the $38 billion: the building, which Microsoft now depreciates over 25 years, the substations and transmission, the turbines ordered for 2031, the fabs and packaging lines. Those outlast any correction. The durable legacy of an AI bust would be power and fabs, plus a fire sale of compute.

The consolation, if you build on this stuff rather than finance it, is that intelligence gets cheaper in every branch. In the boom branch, tokens get cheaper on schedule: chip performance per dollar compounds at 49 percent a year and each hardware generation brings a two- to ten-fold drop in cost per token. In the bust branch, they get cheaper faster: stranded GPUs, stranded power, labs selling capacity at marginal cost to service debt. The bubble question decides who ends up owning the capacity. The capacity exists either way. Plan for abundance.

What I’m watching is the ratio. Capex is a bet that volume outruns price for long enough for revenue to catch it. The back of my envelope:

Mid-2026 Growth rate
Identifiable AI revenue, run-rate (OpenAI, Anthropic, Microsoft AI, Amazon AI, with overlaps) ~$150-200B 2-3x per year
Big-four hyperscaler capex, 2026 ~$750B ~1.7x per year
Required end-customer revenue by Sequoia’s rule of thumb (4x Nvidia’s run-rate) ~$1.2T

At those slopes the gap closes from four- or five-to-one now to parity around 2029, which is where the more careful forecasters put it too; Goldman doesn’t see AI supply and demand balancing until the first half of 2028. That makes the next two years the dangerous window, when capex is still compounding and revenue hasn’t caught it. The signals, in rough order of how early they’d move: OpenAI’s next funding round; Oracle’s credit spreads; whether Stargate’s remaining sites energize on their late-2028 dates; one-year H100 contract prices, the single best anti-bust datapoint right now; and token volume, which has to keep growing faster than the commodity price falls. The day it doesn’t, revenue stops growing and the whole tower is built on a flat line.

I don’t think that day is close. I do think the middle tier of the capital stack is going to have a very bad year before 2030, and that when it does, the GPUs will change hands at a discount, the power will still be there, and the people who bought the discount will run the deployment period on it.

I’ve been on the wrong side of that timing once. The startup I left Microsoft for in March 2000 was Wildseed, built on the idea that phone software was being written by radio engineers, that the radio stack and the operating system would separate into layers the way PC hardware and Windows had, and that the app explosion of the PC era was about to happen on phones. We were describing the mobile revolution seven years before the iPhone and eight before Android, and the wireless industry was blind to it. We were also a software company trying to ship hardware into the capital drought that followed the crash, which is a bad combination. We built the phone in Korea and sold it on second-tier carriers; it was genuinely cool, it was not a market success, and AOL bought the company in 2005. The thesis was right. The deployment period arrived on its own schedule, on networks the telecom bubble had overbuilt and the crash had repriced, and it arrived for other people. That’s the shape I expect here. The last GPU buildout I was part of, in 1998, was financed by teenagers buying Quake. This one is financed by Blue Owl. The silicon doesn’t care who was right first.

Zero times infinity isn’t a number. It’s whatever the two curves happen to be doing when you look, and right now the price is heading to zero, the volume is heading to infinity, and the product is a bill. It’s being paid by Nvidia’s shareholders, Oracle’s bondholders, the Gulf sovereign funds, the pension plans behind the private credit, and everyone whose laptop cost 17 percent more this year. It comes due for them, not for you. And when it does, the zero gets a little closer.


References: