
The AI industry has spent the last two years treating scale as a mainly physical challenge: more chips, more power, more data centres, more capital. That view is understandable, but it leaves out something more fundamental. Data center energy demand is set to more than double from 415 TWh in 2024 to roughly 945 TWh by 2030, yet $156 billion in data center projects was blocked in 2025 alone not from lack of funding, but from inability to connect to the grid.
The pressure point increasingly shows up in utilisation, scheduling, thermal headroom, and load management rather than raw capacity. GPU clusters run at 30–60% utilisation across the enterprise, and cooling systems consume 40–60% of operating energy using rule-based controls designed decades ago. The gap between what gets built and what can actually be used is not a supply problem. It is a coordination problem.
Why the power story is incomplete
The usual story around AI infrastructure is simple: demand is outrunning supply. While that framing captures part of the picture, it misses a more actionable insight. Research from Duke University’s Nicholas Institute shows that electricity providers can meet data center demand on 350 out of 365 days per year (PDF), the constraint is 15 peak days where coordination breaks down, not a fundamental supply deficit.
Advanced economy grids run at around 30% average utilisation, leaving capacity that exists in theory but is hard to access in practice. In the US, interconnection queues have swelled to 2,600 GW, twice the country’s total installed capacity, with median wait times approaching five years and up to seven years in Northern Virginia. The challenge is not to build more but to unlock what is already there. That is where software becomes decisive.
Where the waste happens
AI infrastructure rarely converts installed capacity into useful output efficiently. GPU clusters operate at 30–60% utilisation in a typical enterprise environment. Between 2022 and 2024, inference costs fell by a factor of 280x driven almost entirely by software improvements, not hardware advances. Google doubled its data center energy efficiency primarily through software optimisation, without adding a single new power plant or chip.
The biggest gains therefore come not from adding more, but from better orchestration of what already exists. That means routing workloads more intelligently, reducing inference overhead, improving thermal efficiency, and adjusting power draw in real time.
The important part is how gains compound across layers. Better workload management raises effective GPU utilisation, which reduces peak load, which improves thermal performance, which frees up more compute headroom. A 10% gain at each of four layers delivers more value than a 40% gain at one. The same infrastructure can do significantly more useful work if the layers around it are better coordinated.
A four-layer view

A useful way to map the system is through four interdependent software layers, together representing a $251–362B market by 2030: grid efficiency ($12–24B), facility efficiency ($4–8B), compute efficiency ($60–100B in recoverable capex terms), and software efficiency ($175–230B). Grid software manages interconnection and flexibility; facility software governs cooling and energy use. Compute orchestration decides how workloads are distributed, and model efficiency reduces the compute needed for a given output. Each layer has its own buyer profile, business model, and competitive dynamics and the interactions between them are where the most defensible companies are being built.
These layers affect one another continuously. A better grid signal can improve cooling decisions, which can improve thermal headroom, which can improve GPU utilisation, which lowers the marginal cost of inference. That is where the real leverage sits now: in the handoff between layers that used to be managed separately. The most valuable companies will be those that sit at the control point where these tradeoffs are made simultaneously.
That is also what makes this moment distinctive. Software efficiency investment accelerated six-fold in 2025. NVIDIA made four acquisitions in 12 months: Run:ai, OctoAI, Deci AI, and CentML, all targeting the software layer above its own hardware. A cooling system, a workload scheduler, and a power signal are no longer three unrelated pieces of infrastructure, they are part of one operating environment, and the market is beginning to price that accordingly.
Why this matters for AI professionals
For researchers and engineers, efficiency has become part of the core infrastructure problem that shapes what can be deployed, where it can be deployed, and at what cost. The 280x reduction in inference costs since 2022 was a software and algorithmic achievement, which means the same lever remains available at every layer of the stack. Systems thinking now matters as much as raw model performance.
For product teams, the question now includes how to deliver performance reliably under infrastructure constraints that vary by location, time of day, and grid conditions. A model that performs well in tests but proves expensive or fragile in production will not scale in the environments where demand is rising fastest. The most commercially durable AI products will be those engineered with infrastructure efficiency as a first-order constraint, not an afterthought.
For platform teams, the value is migrating toward systems that can coordinate power, cooling, and compute dynamically. Stranded capacity utilisation software, that locates and activates unused power and thermal headroom inside existing facilities, is almost entirely unfunded today despite the problem compounding every time a data center hits its power ceiling. The strongest architecture is the one that turns more of the installed base into usable output, not the one that simply adds more hardware.
From isolated tools to connected systems
AI infrastructure is still largely assembled as a set of separate products: one tool for cooling, another for workload routing, another for power flexibility, another for model efficiency. Each solves a real problem, but the aggregate gains are limited when the tools are not designed to work together.
The more interesting opportunity sits at the seams. A platform that reads power availability, thermal constraints, and compute demand simultaneously can influence how the whole system behaves, not just one component. That creates a categorically different kind of value: not a feature advantage, but a structural position inside the operating flow of the facility. Inference optimization players like FriendliAI (inventor of continuous batching), Tensormesh (backed by NVIDIA, AMD, and CoreWeave), and OpenRouter (25 trillion tokens per week) each hold this kind of embedded position in their respective layer.
That is also what makes these systems harder to displace. Once software starts shaping scheduling, cooling, and resource allocation together, it becomes embedded in how the infrastructure runs day to day. The strongest companies in this space are the ones moving from solving one problem to helping run the environment itself and the exit record confirms the acquirer appetite: MosaicML to Databricks for $1.3B, Run:ai to NVIDIA, which made four acquisitions in 12 months.
The new bottleneck
The result is a shift in how AI infrastructure should be understood. The most valuable systems will be those that reduce wasted capacity across layers, particularly when demand spikes and flexibility matters most. Grid orchestration software that turns data centers into active grid participants, responding to price signals in real time, has attracted its first large institutional capital only recently, despite representing the highest-value sub-segment in grid efficiency. The categories where the gap between demonstrated impact and current capital deployment is widest are exactly the ones where the exit market is least developed: grid orchestration, stranded capacity utilisation, and inference optimisation.
As model deployment becomes more operationally intensive, coordination becomes part of the product, not just part of the backend. Once hardware, power, and cooling are expensive and slow to expand, the ability to manage them together becomes decisive. Software increasingly determines whether the system works well, works poorly, or gets stuck altogether making it a critical investment opportunity.



