The end of AI is “high quality electricity”! Under the 945 terawatt-hour demand boom, “power system stability” is taking over the top priority of AI infrastructure

Zhitongcaijing · 2d ago

The huge demand for electricity in heavyweight AI training/inference clusters has long been widely known. What is less well understood, however, is how rapid fluctuations in data center electricity demand can damage critical equipment in these huge infrastructures.

In artificial intelligence infrastructure, where construction is in full swing, high-capacity batteries, generators, chillers, and other critical systems are under so much pressure that they fail frequently or reach the end of their useful life prematurely.

As the artificial intelligence boom accelerates, these technical issues mean additional costs and hidden reliability risks of data center power equipment that were not previously anticipated; even a loss of just a few minutes of uptime can hurt data center developers to make money and strong profit data. At the same time, investors and lenders are already uneasy about the hundreds of billions of dollars spent by individual hyperscale cloud computing companies, and the market is increasingly worried that the lowest level hardware infrastructure in these data centers may depreciate much faster than analysts had previously agreed.

These latest developments regarding “AI's drastic fluctuations in electricity demand are damaging its own data centers” seem to indicate that the valuation anchor of the AI infrastructure construction frenzy is further extending from “the number of GPUs/TPUs” to power supply quality, dynamic stability, power chain equipment life, and effective operation time.

The real hidden bottleneck in AI data centers seems to have been upgraded from “whether enough electricity can be obtained” to “whether it can withstand severe power fluctuations on a millisecond scale and a city scale”: hundreds of thousands of AI GPUs or dozens of Nvidia Vera-Rubin and AMD Helios high-power AI racks will repeatedly impact gas turbines, generators, batteries, transformers, UPS, and cooling systems, causing equipment cracking, abnormal wear and tear, and premature scrapping; once a failure causes the computing power infrastructure to shut down and supply is interrupted, the core of the loss is not the cost of replacing parts. Instead, expensive GPU assets cannot continuously generate profits. At the same time, these dynamic loads may also threaten external grid stability through voltage, frequency, and sub-synchronous oscillations. This is why NERC (North American Electricity Reliability Commission) has listed rapid load changes in large data centers as an important reliability risk.

The story “The end of AI is really electricity” is getting more and more popular! Megawatt rack reconstruction data center equipment investment layout

“The end of AI is electricity” is being further evolved into “the end of AI is stable and controllable high-quality electricity” — energy storage, UPS, high-voltage DC, power electronics, microgrids, and advanced cooling equipment will receive higher value, and the valuation of data center projects must also take into account hidden costs that were previously underestimated, such as accelerated equipment depreciation, downtime losses, and reliability transformation.

The ultimate constraint of AI is the ability to deliver high-quality usable electricity to the chip stably, continuously, and with extremely fast dynamic response. Traditional CPU racks have relatively low power density and slow load changes, while rack-level systems such as Blackwell synchronously couple dozens of GPUs, CPUs, NVSwitch and high-speed network cards: the GB200 NVL72 has a full load power of about 120 kilowatts, and the GB300 NVL72 can reach up to 142 kilowatts; when the training task enters collective communication, checkpoint storage, or inference batch switching, the entire GPU simultaneously increases and decreases power at the millisecond level, creating huge surges, harmonics, and heat loads.

Nvidia expects to start supporting IT racks of 1 MW and above in 2027. Traditional 54-volt DC power distribution will meet the physical limits of copper bay size, transmission loss, and conversion levels. Therefore, every time an AI cluster increases its computing power, it must not only purchase more GPUs, but also simultaneously expand medium voltage transformers, switch cabinets, busbars, rectifiers, UPS, battery energy storage, supercapacitors, backup generators, and coolant distribution units; power conversion, distribution stability, and cooling systems are being upgraded from auxiliary facilities to core production materials that determine whether GPUs can be started.

Global companies' demand for strong rack-level AI computing power from Vera Rubin and AMD Helios further shows that next-generation competition has moved from “chip-level performance” to “computing-power-network-liquid cooling” full-frame collaborative design. The Vera Rubin NVL72 incorporates a power smoothing mechanism and is equipped with a rack local energy buffer of about six times that of Blackwell Ultra. It uses batteries or capacitors to instantaneously supply power when the load suddenly rises and absorbs the remaining energy during a sudden drop to avoid transmitting GPU power spikes to transformers and grids; the 800 volt DC architecture reduces copper, cable size, and multi-stage AC/DC conversion losses by reducing current at the same power.

AMD Helios also integrates 72 MI455x GPUs, 31TB HBM4, centralized power management busbars, and liquid-cooled manifolds into one rack-scale system. This is why the higher the AI computing power density, the value and technical threshold of power equipment will rise faster. The structural benefits will expand from simple power generation capacity to transformers and switchgear, high-voltage DC power supplies, UPS and energy storage, power quality management, liquid cooling, and microgrids, while data centers that lack dynamic load management may erode return on capital due to accelerated equipment depreciation and shutdown losses.

Wall Street financial giant Goldman Sachs predicts that global data center electricity demand will increase by 220% by 2030 compared to 2023, which is equivalent to adding one of the world's top ten electricity consumers. The IEA (International Energy Agency) predicts that global data center electricity consumption will increase from about 415 terawatt-hours in 2024 to about 945 terawatt-hours in 2030, accounting for an average annual increase of about 15%, accounting for nearly 3% of global electricity consumption. Among them, AI accelerated server electricity consumption will increase by about 30%, and the electricity consumption of US data centers will increase by about 240 terawatt-hours compared to 2024, an increase of about 130%, and contribute nearly half of the new electricity demand in the US by 2030.

image.png

Millisecond power consumption fluctuates, and AI is rapidly draining its own infrastructure

Amber Villegas-Williamson, a senior consultant at the Uptime Institute in the UK, who advises power suppliers and data centers on standards and reliability, said: “Artificial intelligence can indeed generate very abnormal electricity demand. It's like keeping a car engine running at a very high speed, and the engine will wear out faster than maintaining a constant speed.”

Data centers have been around for decades, and they consume a lot of electricity to keep everything from your favorite streaming shows to online grocery orders running smoothly. However, the underlying computing power infrastructure designed for artificial intelligence computing is different because its electricity demand is extremely large and fluctuates much more.

A load equivalent to the electricity consumption of a factory, town, or even a city may suddenly appear or disappear within seconds, causing repeated shocks, making the devices connected to it unbearable.

Shannon Miller, founder and president of Mainspring Energy Inc., which develops microgrid projects for industrial and data center customers, said that a 1-gigawatt data center consumes the same amount of electricity as a city like Boston, and half of the load may be switched on and off every few seconds. Some artificial intelligence parks planned in Texas and the Midwest of the United States are even five times larger than this level, and their average electricity consumption is almost comparable to that of New York City.

image.png

As shown in the chart above, AI chip sales data indicates a surge in data center power demand — the US will account for 64% of new AI electricity consumption from 2022 to 2033. Note: The BloombergNEF model estimates the power requirements implicit in AI chips, including computing, networking, and cooling infrastructure.

When training new models, AI data centers put a particularly heavy strain on the power supply system—a process that causes all graphics processors to start at the same time. Like swarms of bees flying in groups in a digital world, or a swarm of fish suddenly changing direction, hundreds of thousands of GPUs simultaneously increase or decrease power in milliseconds.

Drew Baglino, a former Tesla executive who founded Heron Power Electronics Co., said that the power consumption of artificial intelligence facilities sometimes soars to more than 50% of the design capacity, “so a 1 gigawatt facility may consume 1.5 gigawatts of electricity in a very short moment.” The company is developing equipment to manage power fluctuations to match Nvidia's next generation of higher-energy servers that Nvidia plans to launch in 2027.

Most devices are not designed for such drastic fluctuations in electricity usage. Jon Parrera, CEO of energy storage developer Terraflow Energy, likens it to switching directly from six to one gear while driving a Ferrari. “You can't switch that fast,” he said.

This article is based on interviews with more than 30 electricity experts in the US and Europe, including power generation companies and other power suppliers, data center developers, grid operators, utilities, investors, standard setters, insurance companies, and regulators. Almost all respondents said that the physical stress on these facilities is already very obvious.

A number of interviewees said that the crankshaft of the small natural gas internal combustion engine used to generate electricity in the data center has broken. One of them claims that the gas turbines have cracked at the xAI Colossus computing facility in Memphis, Tennessee. The source said batteries were then installed in the system to help calm power fluctuations and reduce the pressure on the rotating turbines. XAI's parent company SpaceX did not immediately respond to media requests for comment.

Andrew Cunningham, CEO of GeoPura Ltd., said that some much smaller data centers in the UK have also experienced turbine cracks. The company is supplying hydrogen for fuel cells at select sites to smooth the flow of electricity.

Jennifer Scanlon, CEO of UL Solutions Inc., which is responsible for testing and certifying the new technology, said that cracks or wear on the equipment could cause arc flashovers — that is, electric current jumps between conductors — which could damage artificial intelligence chips.

A complete set of equipment such as batteries, capacitors, transformers, and flywheels can help stabilize the flow of electricity. However, several interviewees said that in the race against time to build artificial intelligence computing power, new data centers have not fully adopted these technologies. According to the Uptime Institute and others working with the operator, some of the batteries installed for this purpose have to be replaced after only a few months or even weeks of use due to excessive stress.

Villegas Williamson of the Uptime Institute said that data centers around the world, from the Middle East and Africa to Europe and the US, are experiencing this problem.

Downtime rates erode cash flow, and AI computing power infrastructure reliability is beginning to become a financial risk

These issues have led to project delays or limited operations at some artificial intelligence computing facilities, and reduced revenue as a result.

Chris James, CEO of Joulent Inc., said that additional time was set aside in the engineering plan to ensure that the 2.67 gigawatt artificial intelligence campus in West Texas can achieve the 99.999% reliability required by Microsoft. The company is co-developing this facility with energy giant Chevron. This means that the electricity supply will be delayed from 2027 until 2028.

Jason Hoffman, chief strategy officer for data center construction and operator Switch, said that if key equipment fails prematurely, “the financial consequences are not mainly the cost of replacing a pump, a circuit breaker, or an electrical component, but rather the value lost due to expensive computing power that cannot generate revenue due to downtime.”

Revenue losses due to downtime vary widely, with estimates ranging from thousands of dollars to hundreds of thousands of dollars per minute, depending on the type of facility and the workload being run.

A person involved in financing such facilities said that the data center is built on the assumption that once it is put into operation, it will operate 365 days a year, around the clock. However, in reality, the normal operation rate of some facilities is only close to 80%; if this problem is not solved, investors in some projects may be impacted within the next 12 to 24 months.

Any reliability issues will further exacerbate market concerns about whether hundreds of billions of dollars of investment in artificial intelligence can generate corresponding returns. The depreciation rate of another key piece of data center equipment — the GPU rack itself — has also raised questions about whether the industry can actually achieve its promised profitability.

These reliability issues may also disrupt the wider power system. The vast network of high-voltage transmission lines, transformers, and power plants requires continuous calibration, but this work is becoming more difficult year by year due to aging equipment, rising electricity demand, and extreme weather.

The continuous expansion of intermittent wind and solar power generation often causes drastic fluctuations in power supply over an hour, which itself has become a factor in grid instability. Artificial intelligence data centers may magnify this fluctuation exponentially.

“These loads are extremely dynamic, or fluctuate drastically, and can cause grid instability; if not corrected, they may cause power outages or interruptions in power supply,” said Slimant Roy, a power quality expert and global product leader at Schneider Electric in Nashville, Tennessee. He said it is particularly worrying that data centers may cause subsynchronous oscillations in the flow of electricity, thereby damaging equipment connected to other parts of the grid.

“This is already very worrying for global utilities,” Roy said.

Over the past two years, the highest regulator responsible for setting US electricity reliability standards — the North American Electric Reliability Company — has issued multiple warnings and warnings, saying that data centers are one of the biggest risks to grid stability.

According to a report released in September, the North American Electric Reliability Company assessed more than 33 gigawatts of operating data centers in the US and found that about three-quarters of these load models “are insufficient to reflect the dynamic behavior of data centers.” Earlier this year, the agency issued a rare Level 3 alert requiring large data centers to address these immediate risks and submit responses by August 3.

From “false calculations” to energy storage buffers, is the AI computing power industry chain beginning to completely restructure the power supply architecture?

From upstream to downstream of the industrial chain, the AI industry is aware of these problems and is actively studying solutions.

Dion Harris, Nvidia's senior director of hyperscale infrastructure solutions, said that when developing Blackwell GPUs, Nvidia began working more closely with power experts. Blackwell was first released in 2024 and is now widely deployed in data centers.

“We're not only making chips and processors,” Harris said, “We're also using these products in Nvidia's own data centers. Nvidia is working to make the deployment process smoother, “including not only the data center construction, design, and engineering stages, but also the power supply process.”

Data center users have used techniques to smooth out power fluctuations in AI workloads, such as running assisted computations — essentially invalid mathematical operations unrelated to the training process — to keep GPUs running stably. However, at a time when demand for electricity is surging, this approach has been criticized for wasting electricity.

Martha Simcoe Davis, project manager at the Rocky Mountain National Laboratory, said that last year, the lab near Denver, Colorado set up a test platform on behalf of the US Department of Energy to study how to safely connect artificial intelligence to the power grid. She is responsible for projects related to the Department of Energy's Electricity Office.

The test site is equipped with GPUs and on-site power generation equipment, and developers can use this to determine whether their systems can withstand fluctuations in artificial intelligence load. A power supplier said it will use the facility to test batteries, software and other equipment to calm oscillations between the data center and the power grid that could damage both parties.

Simcoe Davis said, “We now have an opportunity to do this right.”