The Efficiency Nobody Measures
Why PUE is the wrong number for the AI era
Steve Williams · babnews.org
Every ranking of data-center efficiency runs on a single number, and it is quietly the wrong one for the era we are now in. The number is PUE, and understanding why it is misleading, precisely when it seems most authoritative, tells you more about the AI buildout than the ranking itself does.
What the scoreboard says
Start with the numbers everyone cites, from the hyperscalers' 2025 disclosures, because they are real and worth knowing.
On power efficiency, measured by PUE, the leaders bunch tightly. Google runs a fleet-wide PUE around 1.09 to 1.10. Meta is right beside it, near 1.09, in the low-1.1 range. AWS reports a global average around 1.14 to 1.15, with its best sites approaching 1.04 to 1.05. Microsoft trails the group slightly at about 1.16. For context, the industry average has sat near 1.5 to 1.55 for years and has barely moved. So all four hyperscalers are dramatically more efficient than the field, and the spread among them, from Google's 1.09 to Microsoft's 1.16, is real but small.
The ranking reshuffles on other axes. On water, Meta and AWS lead with a water-usage effectiveness around 0.19 to 0.20 liters per kilowatt-hour, roughly a tenth of the industry norm, while Microsoft runs higher near 0.30, and Google, strong on power, is a heavy absolute water user, on the order of several billion gallons a year, because some of its power efficiency comes from evaporative cooling that spends water to save electricity. On carbon, Google has the lowest intensity per unit of compute and leads on 24/7 carbon-free-energy matching, AWS is the largest corporate renewable buyer by volume, and Microsoft carries the most ambitious target, carbon-negative by 2030. So "who is greenest" already depends entirely on which definition you pick.
That alone should make you suspicious of any single ranking. But the deeper problem is with the headline metric itself.
What PUE actually measures
PUE is simple, which is its virtue and its trap. It is the ratio of total facility power to the power that reaches the IT equipment. A PUE of 1.5 means that for every watt delivered to a server, half a watt more is spent on cooling, lighting, power conversion, and overhead. A PUE of 1.1 means only a tenth extra. Lower is better, and driving it down genuinely cut waste over the last two decades, from an industry average near 2.5 in 2007 to about 1.5 today, and to 1.1 at the hyperscale frontier.
Read that definition again and notice what it does not contain. PUE measures how efficiently a building delivers power to the servers. It says nothing, at all, about what the servers do with that power. It is a metric of overhead, not of output. It rewards a facility for not wasting energy on cooling. It is completely blind to whether the energy that reached the chips accomplished anything.
In the world PUE was designed for, that blindness did not much matter. In a data center full of general-purpose CPUs doing routine cloud work, the servers were a relatively fixed, well-utilized load, and the overhead was the interesting variable. Squeezing PUE from 1.5 to 1.1 was the real efficiency lever, so the metric pointed at the right thing.
The AI era broke that assumption.
Why it's the wrong metric now
In an AI data center, the chips are no longer a fixed background load. They are the entire cost, the entire point, and by far the largest consumer of the power. And their efficiency, how much useful computation they produce per watt, varies enormously with chip generation, utilization, workload, and software. That is precisely the variable PUE cannot see.
Consider the failure case directly. A data center can run the most expensive GPUs made, at low utilization, on a previous-generation architecture, executing poorly-optimized workloads, and still post a flawless 1.1 PUE. The metric would applaud it, because the power reached the chips with minimal overhead. Whether those chips sat idle, or ran at a third of capacity, or did the same work a newer chip would do at half the energy, PUE has no way to know and no way to say. It measures that the power arrived. It cannot measure whether the power was deserved.
So in the era where the compute is the whole story, the industry's headline efficiency metric measures everything except the compute. It grades the packaging and ignores the product. A perfect PUE on a data center full of underused or inefficient silicon is an efficient delivery system for waste, and the number would look pristine.
The metric that matters, and doesn't exist
The efficiency question that actually matters in the AI era is useful compute per watt: the amount of real computational work, training throughput, inference served, tokens generated, a facility produces per unit of energy consumed. That is the number that tells you whether a data center is genuinely efficient at its actual job.
There is no standardized, disclosed, cross-company version of that number. Performance-per-watt exists at the chip level in vendor benchmarks, but there is no accepted facility-level or fleet-level metric that captures useful work delivered per watt across a data center, let alone one the hyperscalers report. So the industry optimizes and publishes the efficiency of the building, where it looks excellent and the differences are tiny, while the efficiency that counts, the efficiency of the compute, goes unmeasured, unreported, and unranked.
This is not a small gap. It means the entire public conversation about data-center efficiency is happening about the wrong layer. Everyone argues over a 1.09 versus a 1.16 while the variable that could differ by multiples, how much useful work the watts actually produce, is invisible.
The bigger tell: siting beats the provider
There is one more number that puts the whole provider ranking in perspective. The single largest determinant of a data center's actual environmental footprint is not the operator's PUE. It is where the facility sits. The same workload running in Montreal, on a grid that is roughly 98 percent carbon-free, versus Singapore, at about 38 percent, produces five-to-ten times the difference in real emissions. That swing dwarfs the entire spread between the most and least efficient hyperscaler PUE.
So the efficiency contest that gets all the attention, provider against provider, tenths of a point of PUE, is a rounding error next to the siting decision that nobody puts in a ranking. Which ties back to a point I keep arriving at from every direction: the data-center story is a geography story. Where the load lands determines its cost, its carbon, and its consent, far more than which logo runs it.
The honest bounds
A few caveats, so this doesn't overclaim. PUE is not useless; it remains a genuinely good measure of the thing it measures, facility overhead, and driving it down was real progress that the hyperscalers deserve credit for. The leaders are also legitimately far more efficient than the industry field, so the bunching at the top reflects real achievement, not just marketing. And building a useful-work-per-watt metric is genuinely hard, because "useful work" differs across training, inference, and conventional cloud, and no one has agreed how to normalize it. The absence of the metric is a real problem, not simple negligence.
But the direction is clear. PUE answers a question that mattered most in the CPU era, and the AI era is asking a different one. The numbers are also self-reported, with no standardized audit, so even the scoreboard everyone cites is a set of company claims rather than verified facts.
The point
So when you see hyperscalers ranked by data-center efficiency, ask what the ranking is actually measuring. It is measuring the building, not the compute. In the AI era, the efficient data center is not the one that wastes the least power on cooling. It is the one that turns the most watts into useful work, sited where those watts are cleanest. Neither of those is what PUE captures, and neither is what the rankings report.
You do not remove a bottleneck. You relocate it. The efficiency that matters moved from the building to the chip, the workload, and the location. The metric everyone still quotes never followed it, which means the industry is optimizing, and grading itself on, a number that is increasingly beside the point.