An aerial view of an electrical power substation along the Columbia River in Oregon, circa March 2026. A Stanford University and Together AI study indicates that electrical infrastructure, not chips, is the real bottleneck on AI. (Shutterstock/Hrach Hovhannisyan)
The Grid Doesn’t Care How Smart Your Model Is
New research on“intelligence per watt” suggests the real AI race may be won on energy efficiency, not frontier chips.
For three years, the contest for artificial intelligence (AI) has been scored one way: whoever trains the largest frontier model, on the most advanced chips, in the biggest data centers, wins. That premise drives US export controls, which aim to deny China the hardware to build ever-larger systems. It drives the buildout now straining power grids from Virginia to Texas, where firm electricity, not capital, has become the binding constraint on how much compute comes online. And it rests on an assumption worth examining: that the frontier model in the cloud is where the game is decided.
A study published in November 2025 by researchers at Stanford University and Together AI, an American AI infrastructure firm, gives reason to doubt it. The authors propose a single yardstick for AI efficiency: intelligence per watt, the task accuracy a system delivers for each unit of power it burns. Running more than a million real-world queries across 20-odd compact models and eight types of hardware, they found that models running locally, on the kind of chip in a modern laptop, correctly answered 88.7 percent of single-turn chat and reasoning queries. From 2023 to 2025, intelligence per watt improved 5.3 times, and the share of queries a local model could handle climbed from 23 percent to 71 percent.
That trend does not stand alone. The price of intelligence has been falling just as fast. OpenAI’s flagship model cost $30 per million input tokens when GPT-4 launched in early 2023; a little over a year later, GPT-4o did the same work for $2.50, and a cheaper variant matching the original on most tasks arrived at a fraction of that. The point is not the exact figure. It is the slope. Efficiency is the fast-moving part of this system, and it moves faster than anything built out of steel and copper.
None of this makes the cloud obsolete. The hardest queries still route to frontier models and will for years. But it reframes what the competition is actually over, and that reframing runs straight into the energy question that governs everything else in AI.
Why Energy, Not Chips, Is the Real Scoreboard in the US-China AI Race
The power demand of AI has been treated as an obstacle to clear: grids to expand, gas turbines to queue for, reactors to license, all so that larger models can run in larger halls. That is the logic of the current data-center boom, and it is why interconnection queues in the largest US markets now stretch five to seven years. Compute is downstream of watts, and watts are downstream of infrastructure that takes a decade to build. In a working analysis I have been developing on the global data-center buildout, the recurring finding is that the United States and the European Union (EU) will underdeliver against their own 2030 power plans, while China, coordinating generation and grid through the state, will not. Energy, not chips, is where that race is won or lost.
Intelligence per watt flips the logic. If a large and rising share of demand can be met on efficient hardware close to the user, the marginal query no longer requires another gigawatt of centralized, firmed capacity. Efficiency, how much useful inference a system wrings from each watt, begins to matter as much as scale. And unlike a transmission line, efficiency is a fast variable. It more than quintupled in two years, faster than grids can be expanded.
For anyone who has watched energy become the true ceiling on AI ambition, this is the more interesting frontier. The durable advantage may accrue less to whoever assembles the most frontier compute than to whoever converts energy into useful output most efficiently, a contest decided by chip design, manufacturing, and system architecture rather than by the sheer capital poured into data centers.
What It Does to the Logic of Export Controls on AI Chips
US controls rest on a specific theory of leverage: deny China the most advanced accelerators, and you cap the frontier systems it can build. That theory holds only so long as frontier scale is where capability lives.
Push most everyday demand down to efficient hardware, and the leverage erodes at the margin. Here a distinction matters, and it is one a careful reader should hold onto. The Stanford result concerns small models on consumer-grade silicon, laptops rather than server farms. But China’s parallel effort is to build out domestic inference chips at scale. Huawei’s Ascend line, which by its developers’ own testing runs at roughly 60 percent of an Nvidia H100 on inference, is now the default for a growing share of Chinese deployment despite lagging the frontier. A competitor blocked from the highest-end chips, but strong in efficient and deployable hardware across both tiers, has a plausible path to serving the bulk of real-world AI demand without ever matching the frontier.
This is not an argument that export controls are futile. The chokepoint still bites hardest exactly where it is meant to: frontier training, the largest models, the most advanced nodes. It is an argument that export controls target one layer of the stack while the efficiency layer, where denial works poorly and a manufacturing-heavy competitor is comparatively strong, goes largely unaddressed.
The Metric Shapes the Strategy for AI Efficiency and National Competitiveness
There is a reason intelligence per watt sits at the edge of the strategic conversation. It is hard to count. Analysts and officials benchmark AI dominance by what they can tally: chips shipped, parameters trained, data-center megawatts announced. Watts of useful inference resist that kind of scoreboard.
That is the deeper point, and it echoes an older lesson from energy geopolitics. What a competition measures shapes the strategy pursued to win it. Count barrels, and states chase reserves. Count firm capacity, and they chase grids. A frontier-and-chip scoreboard produces a race to build the largest systems and deny rivals the means to match them. A watts scoreboard would reward grid quality, efficient silicon, and the industrial base to put that silicon into wide use, a different race, and one in which the current US approach is not obviously ahead.
A scoreboard change of that kind would carry into policy. It would mean pairing chip export controls with what they neglect: treating grid capacity, transmission, and firm power as instruments of AI strategy rather than permitting problems to be solved later, and measuring progress in useful inference delivered per watt, not accelerators shipped. Denial buys time at the frontier. It does not build the power system that decides who can actually use AI at scale, and that system is the one Washington has left largely to chance.
The Stanford and Together AI paper is a single study, and its authors are careful about its limits: single-turn queries, an idealized routing assumption, frontier models still doing the hardest work. It does not settle how AI power should be counted. But it puts the question on the table, and for a policy architecture built almost entirely on the opposite assumption, that is worth reckoning with before the next round of controls is drafted.
About the Author: Fyodor Dmitrenko
Fyodor Dmitrenko is a geopolitical analyst and researcher specializing in sustainable development, energy policy, and governance on the Eurasian continent. He is affiliated with Sciences Po Paris, where he conducts research under the supervision of Professor Tatiana Mitrova. He has contributed articles on developments in energy markets and international trade flows at Reuters News Agency’s CIS office and for emerging think tanks such as India’s TheGeostrata and the Paris section of the French-MFA-affiliated Andalus Committee, which deals with EU-global south relations. He has also engaged with leaders in the sustainable development field at the Guiyang Ecological Forum as a Sciences Po delegate to the Tsinghua Global Youth Dialogue, and interviewed policymakers such as former Brazilian central bank head Gustavo Franco and former Swiss President Simonetta Sommaruga as the Sciences Po delegate to the Warwick Economic Summit.
The post The Grid Doesn’t Care How Smart Your Model Is appeared first on The National Interest.