
Battery energy storage systems (BESS) are growing faster than the teams managing them. In 2020, Hornsdale Power Reserve was the largest lithium-ion battery installation in the world at 194 MWh. In 2025, the Edwards & Sanborn solar and BESS project clocked in at 3.2 GWh. That’s more than 16 times the capacity, in just five years.
Much of that growth traces back to the volatility renewables bring to the grid. Solar and wind push prices around, create curtailment, and open the arbitrage spreads that batteries exist to capture. Rising data center demand is adding to the pull on new generation, but the core driver is the intermittency itself.
Whatever the energy mix, the batteries being installed are not there only as insurance against downtime. They are commercial assets, measured in real time against contracts that may already be out of date. More capacity, built faster, means more sites per manager, and more places for a silent data error to hide.
The expertise needed to manage systems at this scale is scarce, and AI is being offered as a tool to help stretched portfolio managers keep on top of both maintenance and markets. But the most useful thing AI can do is not produce more analysis — it is to reduce the distance between a developing problem and the person who needs to know about it. And it only works if your data is accurate in the first place.
The challenge lies in cleaning the data, not analysing it
State of charge (SOC) inaccuracy is one of the most consequential yet underappreciated problems in monetizing a battery asset. It sounds simple. If you don’t know how much energy is in the battery, you don’t know how much you can sell back to the grid.
Before any analytics platform can tell you something reliable about SOC, the underlying data must be made trustworthy. Modelling is hard in its own right. But a good model on bad data still returns a wrong answer: just a more convincing one. Data quality is the part that gets the least attention and does the most damage when it is missing.
This is where most analytical failures start. One unreliable input corrupts everything downstream, and the effect compounds across an entire portfolio.
A miscalibrated sensor, an energy management system (EMS) with inconsistent timestamping, or a communication dropout that fills with zeros rather than nulls could all produce analysis that looks authoritative but isn’t. When an operator is managing dozens of sites, each generating thousands of signals at sub-minute intervals, opportunities for these failures multiply.
The commercial cost is real. My team worked recently on a 75 MW site on the Texas Interconnection grid (ERCOT) where the battery’s SOC was being underreported. The system did not recognize the energy it had available to sell. 3.8 MWh of tradable capacity sat idle during an overnight price spike, and the same pattern repeated most days. Across a portfolio of sites doing the same thing, that adds up to a six-figure annual loss. The fix required no new hardware; only a software-based SOC correction through the existing EMS. But it would never have been identified without the data quality work that made the problem visible in the first place.
Why the depth matters
There is always an instinct to bring a capability like this in-house. With battery analytics, that instinct underestimates the problem. Building good models is hard, and it stays hard. But a model is only ever as good as the data underneath it, and that data is the part you cannot shortcut. It must reflect how batteries actually behave across chemistries, vintages, climates, dispatch regimes, and operational histories, and most of its value sits in the failures, the rare and expensive things that go wrong. That kind of dataset takes years to accumulate.
Those failure cases are the ones that matter most commercially, and they are exactly what a model trained on a narrow dataset has never seen. Which is why the model cannot be purely statistical. You need deterministic parts, the physics and the contractual limits you can encode directly, constraining the probabilistic parts that learn from data. Get that balance wrong, or starve the data side, and the system does not just miss things; it produces confident, plausible readings that are wrong, and false negatives for faults it has never been trained on.
The underused KPIs
Some of the most valuable performance indicators in battery storage are also the most consistently overlooked. Setpoint deviation tracking, round-trip efficiency analysis, and availability monitoring all fall into this category. Every operator knows they matter. Measuring them consistently and correctly is harder than it looks, and the cost of getting it wrong shows up directly in missed or lost revenue.
Availability is the percentage of time an asset is operational and ready to perform. It sounds like the simplest metric of the three until you look at how most operators actually define it. A battery can count as ‘available’ while operating below the capacity threshold its contract requires. That single overstated figure can cause you to miss dispatch obligations, triggering penalties or breaching contracts.
Setpoint deviation is another good example. A setpoint is an instruction sent to a battery asset, telling it how much power to charge or discharge at a given moment. Setpoint deviation is the gap between that instruction and what the asset actually delivers. On any individual site that gap can look small enough to ignore. But aggregated across a portfolio, over time, and mapped against live market conditions, it becomes a material drag on revenue that no single site operator has the visibility to catch.
Round-trip efficiency is the measure of how much energy you get back out of a battery relative to what you put in. On any individual report it looks like a straightforward number. It is usually calculated far more crudely than that number implies, ignoring temperature dependence, partial cycles, and where exactly the measurement boundary sits. An efficiency figure that looks acceptable on paper can hide losses that compound over time.
AI can close these gaps, but only when it is working from data that is clean enough, and a model trained deeply enough, to tell signal from noise, especially at the margins where the commercial value sits.
Efficient decisions, not more decisions
An alert that arrives with a root-cause assessment, a severity ranking, and a clear indication of the commercial consequences is a fundamentally different tool from a threshold alarm. The latter requires the operator to start an investigation, while the former gives them a decision. In a sector where portfolios are growing faster than teams, dispatch windows are tight, and the commercial stakes of a missed fault are high; that difference matters enormously.
The operators and asset managers running these portfolios are not looking to be replaced by AI. They are looking for tools that make their judgement more effective, grounded in the experience of real assets and real outcomes. The data is what makes that possible.

