AI & Technology

The AI battery race will be won by better data, the process to generate it – and the discovery engine it enables

By Mathias Ingvarsson, Founder and CEO, Holyvolt Group and Executive Chairman, Wildcat Discovery Technologies

Any model is only as good as the data behind it, and most battery R&D never generates that data at the scale or structure AI demands which is why claims of AI-driven breakthroughs have so far outpaced demonstrated results. The answer is not a better algorithm. We believe the key to unlocking the full potential of AI in battery development is high-throughput experimentation, generatingstandardized, high quality, granular data sets at scale which continuously feeds the AI discovery engine. 

Thirty-five years is a very long time in R&D terms but since the commercialisation of Lithium-Ion battery technology in 1991, the state-of-the-art has remained fundamentally unchanged. The global battery industry has achieved huge reductions in cost and significant increases in energy density since then, but further improvements are still needed, just as they are in all performance attributes – including safety, cycle life, and sustainability. Without these we cannot successfully transition from fossil fuels to renewable energy. 

The massive challenges in making a breakthrough are further compounded by the fact that China has long been the world leader in battery technology, and dominates the supply of raw materials and finished cells. So how does the West compete? AI has been hailed by some as the solution, but the reality is, of course, far more complex than simply taking what’s available, feeding it into a model, and expecting to get a next-generation chemistry that’s ready for mass production. 

Few such efforts have produced real results so far, and the reasons are rarely shortcomings in algorithms or computing power. Successive reviews of machine learning in battery materials reach the same conclusion: the data are scarce, heterogeneous, and drawn from too few comparable samples for models to generalise dependably. Predictive accuracy follows the quality of that data far more than the sophistication of the algorithm.1

We shouldn’t be surprised. A model trained on public-domain data inherits the available literature’s deepest bias: journals and white papers focus on what worked. Experiments that failed or fell short – negative data points – tend not to be published. Much of the work done within the battery supply chain remains strictly confidential. In other words, not enough data, and not enough of the right kind. For an AI model this is crippling. 

A platform to build on 

For these reasons, AI alone cannot deliver a meaningful increase in research speed or efficiency without high throughput testing to correlate and confirm its outputs. Already, high throughput parallelises what had been a sequential process, reducing development timescales by a factor of ten. Intelligently combining the two will enable true AI-led development, with the potential to compress timings by a further order of magnitude. 

A promising application in the near-term is simplifying design of experiments: in 2020, researchers used machine learning to identify charging protocols from a total of 224 candidates in only 16 days instead of 500.2 Cutting that down to just five or less using the technology available now is feasible – but only if you have the right dataset. 

Wildcat does, and has spent more than 18 years compiling a standardised, 500TB dataset without equal, built around 20,000 parallel test channels, more than 500,000 cells, and 75 million cycles. Everyexperiment has been recorded exactly the same way, in the same laboratory, and with failures valued as much as successes. 

Of course, that dataset is continuously growing, with more than 80,000 cycles added daily. A recent addition captured gas generation in pouch cells across thousands of standardised tests, something most programmes never measure at this scale or consistency. In one client project the objective was to fail fast: many variants of a promising electrolyte were screened and ruled out in months, not the years a conventional programme would have taken to reach the same result. 

From measurement to prediction 

As that dataset deepens, new use-cases become possible. One of the most valuable is accurately predicting cycle life: when a dataset is generated for this purpose, models can forecast long-term performance from the first handful of cycles instead of waiting months and thousands of cycles for a cell to age6 – but only when the dataset has been deliberately designed for that purpose. Such models are also narrow – tied closely to the chemistry they were trained on, so a material change means retraining – and that is only feasible with enough structured data to retrain quickly. 

The same logic applies to some of the battery sector’s bolder claims. Computational screening can propose promising new compositions, but a prediction is only a hypothesis until a material is actually made and tested. High throughput closes that loop because it enables validation at scale. 

AI’s extraordinary potential cannot be realised without either the structured, high-quality data it needs to learn from, or the means of validating outputs within a timescale short enough to ensure they remain relevant. Wildcat has both. And as we strengthen our collaborative R&D programmes with stakeholders throughout the battery sector, and with new partners from the AI world, we’ll all progress much faster on the chemistries that matter most. 

Regaining the advantage 

The need is urgent. The IEA expects battery demand to roughly triple by 2030,7 but with most of the world’s cell manufacturing capacity located in China, the West cannot compete on commodity production. It can only compete by discovering better chemistries and commercialising them faster – which is precisely what the approaches described here make possible. 

AI itself is one domain where the West holds a commanding lead, and is at the frontier of foundation model development, compute infrastructure, and AI research talent. That advantage is meaningful if it can be directed at the right problems – and battery chemistry, alongside materials discovery more broadly, is exactly the kind of problem it was made for. The models are ready. What has been missing is the industrial method to feed them and to validate what they produce.  

The combination of AI, high quality data, and high-throughput experimentation will move battery discovery from iteration to prediction – and ultimately to a discovery engine that not only designs the next experiment but finds and then proves the next breakthrough material. The advantage does not arrive as a single leap in performance. It compounds and learns with every cycle run and every result fed back into the system that generated it, and this is the transformation that will deliver the clean energy technologies of tomorrow. 

References 

1 H. Xue et al., ‘Solutions for Lithium Battery Materials Data Issues in Machine Learning: Overview and Future Outlook,’ Advanced Science, 2024. doi.org/10.1002/advs.202410065 

2 P. M. Attia et al.; ‘Closed-loop optimization of fast-charging protocols for batteries with machine learning,’ Nature, 2020 https://pubmed.ncbi.nlm.nih.gov/32076218/ 

5   P. Raccuglia et al., “Machine-learning-assisted materials discovery using failed experiments,” Nature 533, 73–76, 2016. nature.com/articles/nature17439 

6   K. A. Severson et al., “Data-driven prediction of battery cycle life before capacity degradation,” Nature Energy 4, 383–391, 2019. nature.com/articles/s41560-019-0356-8 

7  IEA, Global EV Outlook 2025: Electric vehicle batteries, Paris, 2025. iea.org 

Author

Related Articles

Back to top button