DataAI & Technology

Half of 2026’s AI Data Centre Pipeline May Not Arrive. The Industry Needs Another Architecture

By Daniil and David Liberman, co-creators of the decentralised AI compute network Gonka

For the past several years, the AI industry has spoken as though its primary constraint were access to advanced chips. In 2026, however, the more immediate constraint is becoming something older and less glamorous: electricity, transformers, substations, transmission lines, permits and public acceptance.

At least 16 gigawatts of large data centre capacity was scheduled to come online globally in 2026 across roughly 140 projects. Yet only about 5 gigawatts were under construction earlier this year. Analysts have begun to conclude that 30% to 50% of the planned pipeline may not become operational before year-end.

Some of those projects will eventually be completed. But the distinction between cancellation and a multiyear delay matters less than it appears, because AI demand is growing now, chips depreciate quickly and model architectures change faster than power infrastructure can be approved.

The grid does not read press releases. An announced data centre is a financial intention; a powered rack is computational reality.

The Power Emergency Is No Longer Theoretical

During the July 4 weekend, the US Department of Energy extended an emergency order under Section 202(c) of the Federal Power Act. The order authorised PJM, America’s largest regional grid operator, to direct backup generation resources at data centres and other major facilities to operate as a last resort during a severe grid emergency.

This followed similar interventions involving backup generation at large facilities in January and May. Data centres did not single-handedly cause these emergencies, but the orders show that data centres  are no longer ordinary commercial customers sitting quietly at the edge of the electricity system.

They are becoming power-system actors in their own right. A single AI campus can request hundreds of megawatts or even more than a gigawatt, which is comparable to the demand of a city and cannot be added to a regional grid as casually as another office building.

The US Department of Energy estimates that data centres consumed approximately 4.4% of US electricity in 2023 and could consume between 6.7% and 12% by 2028. The International Energy Agency expects global data centre electricity consumption to more than double by 2030, reaching around 945 terawatt-hours, with AI as the most important driver of that growth.

The Next Bottlenecks Are Institutional

The first bottleneck is electricity supply. The next bottlenecks are likely to be the institutions responsible for allocating it.

Grid operators must decide which proposed loads are real, who should finance new generation and transmission, and whether households and existing industries should carry part of the cost. The Federal Energy Regulatory Commission (FERC) has already warned that speculative and duplicated data centre applications can clog interconnection processes, distort demand forecasts and create unnecessary costs for consumers.

Local opposition will also become harder to dismiss. Communities are being asked to accept new transmission infrastructure, water consumption, backup generators and changes in electricity prices, often in exchange for fewer permanent jobs than a similarly sized industrial development would provide.

Then there is the financial problem. A company committing tens of billions of dollars to a purpose-built campus must recover that investment over many years, even if a new chip, inference technique or model architecture makes the facility less competitive before its depreciation schedule is complete.

AI develops in software time. Energy infrastructure develops in civil-engineering time.

Training Architecture Is the Wrong Default

Large, tightly integrated clusters are valuable. Frontier model training often requires thousands of accelerators communicating over extremely fast networks, and this workload may justify purpose-built campuses with dedicated power.

But the industry has begun treating the architecture required for the most demanding training runs as the default architecture for nearly every AI workload. That is an expensive assumption.

Most people and businesses only meet AI after training is finished. Inference happens when a trained model answers a question, writes code, analyses an image, operates an agent or processes a business workflow.

These requests do not all have to be executed inside the same building. A model can be replicated across many clusters, while individual requests are routed according to latency, capacity, hardware, energy availability and price.

Inference Changes the Economics of Compute

Training resembles a major construction project: concentrated, capital-intensive and completed in large campaigns. Inference resembles a utility: continuous, geographically distributed and highly sensitive to utilisation.

This distinction changes the relevant economic question. The industry should keep asking, “How do we build the next gigawatt campus?” And it should also ask, “How do we obtain more useful AI output from the GPUs and electricity that already exist?”

Modern inference systems frequently separate hardware into pools based on models, latency requirements or customer classes. Microsoft researchers have found that this type of siloing can leave expensive accelerators underutilised when demand varies across workloads and regions.

A broader compute market can match workloads with appropriate hardware instead of forcing every request onto the newest accelerator in the most expensive location. Latency-sensitive requests can stay close to users, while batch processing, evaluation, synthetic data generation and other flexible workloads can move to locations where capacity and electricity are available.

Recent research increasingly treats AI inference as geographically assignable electricity demand. Requests can be routed among multiple computing locations while respecting latency, capacity, energy-price and emissions constraints.

Not every computer can run every model; memory, bandwidth, networking, reliability and security requirements remain real. The argument is narrower: co-location makes sense as a technical requirement where necessary; as an economic doctrine applied to the entire AI market, it deserves scrutiny.

Whoever Controls Compute Controls the Market

AI models are information. Once trained, their parameters can be copied at relatively low cost.

The scarce layer is the physical infrastructure required to run them. That includes chips, electricity, cooling, networking and access to the facilities where those resources are assembled.

Control over that layer enables charging rent across the AI economy. A business may own its application, customer relationships and proprietary data, but if every request must pass through one of a small number of compute providers, it remains dependent on their pricing, capacity allocation and terms of access.

Malice plays no part in this. A market in which only a few organisations can finance and energise the required infrastructure produces such dependence predictably.

The risk is that AI reproduces the cloud market’s concentration at a much more consequential layer. Compute will not simply host software; it will increasingly perform the reasoning, research, customer service, programming and decision-making inside that software.

Emerging Economies Have a Third Option

This concentration creates a particularly difficult choice for emerging economies. They can spend billions attempting to reproduce infrastructure already being built at much greater scale in the US and China, or they can become regular  customers of foreign providers.

The World Bank notes that high upfront costs make frontier AI chips and large data centres prohibitive for many low-income and middle-income countries. As of June 2025, high-income countries accounted for 77% of global data centre capacity, while low-income countries accounted for less than 0.1%.

Trying to outbuild the US or China one country at a time is unlikely to close that gap. A regional economy cannot achieve sovereignty merely by constructing a smaller version of infrastructure controlled elsewhere.

There is a third option: shared and interoperable compute. Countries, universities, telecommunications companies, energy producers and independent infrastructure operators can contribute capacity to larger markets rather than attempting to own the entire stack individually.

A country may not have enough resources to build the world’s largest AI campus. It may still possess competitively priced renewable energy, existing high-performance computing facilities, engineering talent, underused industrial sites or regional connectivity.

Connecting these assets creates a different type of sovereignty. It rests on having multiple routes to computation, portable workloads, open models and the ability to switch providers without rebuilding an entire application. Owning every chip is beside the point.

Policy Can Build Markets Alongside Monuments

So what should the policy response be? Hardly to stop building data centres altogether. But a handful of enormous campuses is clearly not the only credible path to AI capacity.

Governments can accelerate interconnections for flexible computing loads that agree to reduce consumption during grid stress. They can require transparent readiness standards so speculative projects do not occupy interconnection queues, while ensuring that developers finance the infrastructure required to serve them.

Public procurement should favour interoperability and workload portability rather than permanent dependence on a single vendor. Research funding should support open models, efficient inference, heterogeneous hardware, secure execution and systems capable of routing workloads across multiple providers.

Energy policy and AI policy can no longer be written separately. Every national AI plan now contains an implicit electricity plan, whether policymakers acknowledge it or not.

The Future Will Be Hybrid

Centralised and distributed infrastructure are not mutually exclusive. Some training workloads will continue to require hyperscale facilities, while inference, evaluation and many enterprise workloads can be served by federated pools of independent clusters.

The strongest architecture is therefore likely to be hybrid. It will combine large campuses where physical proximity is essential with open compute markets wherever requests can be distributed.

Suppose half of the 2026 pipeline does miss its scheduled delivery. Would that prove AI demand has been exaggerated? The more likely lesson is that the industry has relied too heavily on a single method of supply.

We should build new generation, modernise grids and improve data centre efficiency. But before asking every community to accommodate another gigawatt-scale facility, we should also make better use of the computing capacity already deployed worldwide.

One organisation building the largest possible machine will not get us there. AI abundance arrives whenmany machines, owned by many participants, can compete to serve the same request.

Speakers Bio:

David and Daniil Liberman are serial founders and operators who have spent two decades building across AI, media, fintech, and augmented reality as a sibling team. They are the co-creators of Gonka, a decentralized network for high-efficiency AI compute that directs nearly 100% of GPU cycles toward productive AI inference and training rather than the wasted computation of conventional blockchains.

Author

Related Articles

Back to top button