Chapter 9
The Room to Run
Aa
Energy limits speed. Entropy limits memory.
Physics permits computation far more efficient than anything we have built, though comparisons between an irreversible bit reset, a switching event, a floating-point operation, and an entire system cannot be collapsed into one distance from a floor. Seth Lloyd's ultimate limits lie far beyond current machinery.1 The economically relevant horizon is set by engineering constraints that bind long before ultimate physics, and by financing and delivery constraints that can bind before engineering. The question is whether technical headroom becomes cheaper useful computation or merely more demanding capability.
Biology supplies an existence proof, not a conversion table. A human brain sustains cognition on roughly twenty watts, while estimates of its computational work vary by orders of magnitude and do not map cleanly onto floating-point operations.2 That modest power envelope shows that useful cognition need not require the energy profile of present machines. It does not establish a numerical efficiency gap that silicon must close, still less a schedule for closing it.
The engineering record is more legible when kept within a common measure. Koomey and his coauthors found that computations per kilowatt-hour roughly doubled every 1.57 years across the systems they studied through 2009.3 Prieto and coauthors report a slower 2.29-year doubling time in their later TOP500 and Green500 sample.4 The samples, periods, and workloads differ. Together they establish a long improvement in measured energy efficiency and a later slowdown within one important class of high-performance machines, not a universal law or a single causal account of the change.
The data-movement wall
The distance between current systems and the Landauer floor is misleading as a measure of usable headroom. Mark Horowitz's 2014 data for a 45-nanometer process put a 32-bit floating-point multiply at about 3.7 picojoules and an off-chip DRAM access at about 640 picojoules.5 Within that example, fetching an operand cost roughly 170 times as much as multiplying it. The ratio is process- and workload-specific, but its lesson survives: arithmetic can become cheap while moving data through a hierarchy remains dear. A physical limit on bit erasure tells us little about which level of a deployed system will consume the next watt.
The economic question is whether headroom becomes abundance. It does not, at least not automatically. Headroom describes the distance between current practice and physical possibility. It does not describe the relationship between efficiency gains and demand growth. If efficiency improves by a factor of two and demand grows by a factor of ten, total energy consumption rises fivefold even as cost per operation falls.
Headroom is not abundance. Efficiency can lower the cost of a fixed computation while expanding the range of computations worth buying. Patterson and coauthors' accounting for GPT-3 shows how materially consequential a single large training run can already be.6 It does not establish the size of an undisclosed successor run or a law of ever-rising budgets. It does make the economic question unavoidable: what happens to total demand when each unit becomes cheaper while uses multiply?
Jevons argued that economy in the use of fuel could enlarge its consumption by making more uses profitable.1 The mechanism applies to computation: cheaper inference makes more tasks worth attempting, cheaper training makes more experiments affordable, and greater capability brings more ambitious applications within reach. For an undertaking extending what it can attempt, an efficiency gain can buy a larger experiment instead of a smaller bill. Whether that expansion outweighs the saving per operation depends on how much new work is undertaken.
The physical evidence is already visible. The International Energy Agency's April 2025 special report estimates that data centers consumed approximately 415 terawatt-hours of electricity in 2024, roughly 1.5 percent of global electricity demand, and projects about 945 terawatt-hours in 2030 under its base case.7 That is a forecast, not destiny. Its importance is that efficiency improvement and aggregate demand growth already coexist in the same institutional plans.
ERCOT's large-load interconnection queue supplies a sharper local instance. By December 2025 it contained 238.6 gigawatts of requests, more than 70 percent attributed to data centers, while 7,502 megawatts had been approved to energize since January 2022.8 A queue is an inventory of requests, not a forecast of construction; speculative and duplicate projects can inflate it. Even so, the distance between requested and approved capacity reveals an institutional passage that software cannot shorten by becoming more capable.
Interconnection, transformers, permitting, and capital can therefore bind deployment at a particular site. Elsewhere the binding constraint may be demand, model capability, integration, or the authority to act. No single bottleneck governs the industry. The harder proposition is that computation does not escape physical economy when its algorithms improve. It changes which material and institutional constraints become valuable.
Computation may continue to become cheaper per operation, but efficiency and abundance do not move together automatically. The bits that encode a trained model can be duplicated at low marginal cost, and another path may sometimes reproduce its capability more cheaply than the first. Deployment still requires a material system. Every query, rendered frame, and executed decision draws on hardware and power that must exist somewhere.
This creates a tension that will shape the economics of the emerging regime. Falling computation costs can make some cognitive services more abundant while the means of producing and deploying them remain scarce, unevenly located, and institutionally rationed. The cloud is grounded in concrete, copper, silicon, permission, and finance. Control of that ground can confer advantages that do not reduce to software, though which gate matters will change with the workload and the place.
Source notes
Footnotes
-
William Stanley Jevons, The Coal Question, second edition (Macmillan, 1866), chapter VII, “Of the Economy of Fuel,” especially the distinction between domestic saving and expanding manufactures. Jevons supplies the historical argument. Its application to computation is the chapter’s conditional economic reasoning; no general frequency or magnitude of rebound is established here. ↩