Interview

What it takes to make Healthcare AI work in the real world

Healthcare AI is moving rapidly from research environments into systems that can influence screening, outreach, and clinical decision-making. But as predictive models move closer to real patient care, accuracy alone is no longer enough. The quality of the underlying data, the way clinical outcomes are defined, and the decisions made around deployment can all shape who ultimately benefits from these systems.

Mansi Goel’s work focuses on applying machine learning to large-scale electronic health record data to identify earlier signals of disease and support clinical decision-making. Her experience includes early disease detection across areas such as arrhythmia, liver disease, and Type 1 diabetes, with a particular focus on what happens when predictive models move beyond development and into real-world healthcare.

Working with real-world clinical data has given her a close view of the challenges that emerge long before an AI model reaches production from incomplete and unevenly collected patient histories to choices around target variables, patient cohorts, thresholds, and model evaluation.

We spoke with Mansi about what makes healthcare data uniquely difficult to model, where bias can enter an AI system, how teams should think beyond traditional accuracy metrics, and what it will take to build clinical AI that is not only technically strong, but useful, accountable, and safe in practice.

From Clinical Data to Earlier Detection

What first drew you to healthcare AI and clinical data science?

What first drew me to healthcare AI was the opportunity to use data science to solve problems that have a direct human impact. My background in mathematics and data science gave me a natural interest in finding patterns in complex data, but healthcare added a very different dimension to that work. In healthcare, a pattern is not just an interesting statistical finding, it can influence how we understand risk, when we intervene, and potentially whether a patient gets the right care at the right time.

That became particularly meaningful through my work on early disease detection for cardiac arrhythmia. The idea that we can use information already present in a patient’s clinical record to identify someone who may be at elevated risk, potentially before they develop obvious symptoms or experience a serious event, is powerful. It creates an opportunity to move from a reactive model of healthcare toward earlier identification and intervention. At the same time, it has made me very conscious of the responsibility that comes with making predictions about people’s health.

That combination is what keeps me interested in healthcare AI: the technical challenge of extracting meaningful signals from complex clinical data, and the possibility that, when done responsibly, those signals can translate into earlier action and better outcomes for patients.

What makes electronic health record data so valuable and so difficult to work with?

Electronic health record data is valuable because it captures a patient’s health journey in a way that few other data sources can. It contains clinical observations, diagnoses, medications, procedures, laboratory results, and patterns of healthcare utilization over time. When analyzed appropriately, these signals can help us understand risk and potentially identify problems earlier.

In our work on early disease detection for cardiac arrhythmia, for example, seemingly routine information in a patient’s record can collectively reveal patterns associated with a higher likelihood of future risk – creating an opportunity to look more closely before a serious event occurs.

But that richness is also what makes EHR data difficult to work with. It was primarily created to support patient care and clinical documentation, not to produce clean datasets for research or machine learning. Records can be incomplete, inconsistent, differently documented across providers, and influenced by how and when a patient interacts with the healthcare system.

There is also the challenge of separating a true clinical signal from artifacts of documentation or healthcare utilization. So, before building a sophisticated model, a great deal of care is needed to understand what the data actually represents. For me, that understanding is just as important as the algorithm itself.

How can AI help identify diseases earlier than traditional clinical workflows?

The unique capability of this technology is that it can start to see a pattern of risk before there has been a “triggering” clinical event to respond to. Conventional healthcare generally starts because of a symptom, an abnormal test, or a patient presenting with concern about a particular health issue. AI can be applied to large volumes of clinical data to identify signals that may not seem consequential on their own, but which together indicate a higher future risk.

I see this play out in my own work in early disease detection, including cardiac arrhythmia. A patient may have no prior documented arrhythmia and no overt cardiac symptoms, but upon analysis of their existing clinical data by AI, we can see that they merit closer inspection. Technology is helping us see part of a patient’s clinical journey that may be invisible to us until disease actually becomes clinical.

For clinicians, this shifts the starting point of the healthcare workflow. Instead of healthcare starting when a problem is presented, we can use these systems to identify people who should be proactively monitored or evaluated even before they present with that problem.

Using AI in the Lucem Health–iRhythm partnership, we can identify arrhythmia earlier in patient populations with elevated risk and place them on proactive, targeted cardiac monitoring. The critical change is that AI is not replacing clinical assessment – it is helping ensure that the clinician looks at the right patient at the right time.

For patients, this offers something that conventional healthcare cannot always provide at this stage: a chance to intervene before disease progression brings them into the healthcare system. Preventing that first event is what I think is the real promise of predictive technology.

What Healthcare AI Actually Learns

Which decisions have the biggest impact on what a healthcare AI model ultimately learns?

The decisions that have the biggest impact are often the ones made before the algorithm is selected: what question we ask, what outcome we define, what data we use, and who is represented in that data. These choices determine what the model has an opportunity to learn.

For example, deciding to predict a clinically meaningful outcome versus a proxy – such as healthcare utilization – can lead the model to learn very different patterns. Similarly, decisions about which variables to include, how to handle missing data, and how patients are selected can introduce assumptions and biases that an algorithm may simply reproduce.

Our work on early disease detection for cardiac arrhythmia illustrates this. If our objective is to identify people who are genuinely at elevated arrhythmia risk early enough for clinicians to intervene, then we have to carefully define what “risk” means, the prediction window, and which information would realistically be available before that point. Otherwise, the model may learn signals associated with having received more healthcare rather than signals associated with arrhythmia itself.

So, in my view, the most consequential decisions are not necessarily technical ones. They are the decisions about the question, the data, the population, and the outcome. The algorithm learns from the boundaries we set for it.

Can bias enter a healthcare AI system before the model is even trained?

Absolutely. In healthcare AI, bias can enter long before we choose an algorithm or train a model. It can come from the patients represented in the data, how healthcare is accessed, what clinicians choose to document, how diagnoses are recorded, or even how we define the outcome we want the model to predict.

If certain populations are underrepresented or the data reflects unequal access to care, the model can learn those patterns without anyone intentionally introducing bias.

This is particularly important when building AI models for early disease detection. In our work on arrhythmia, for example, the goal is to identify people who may be at elevated cardiac risk so that there is an opportunity for earlier clinical evaluation. But if the underlying data does not adequately represent the broader patient population, a model could perform differently across groups – and an apparently strong overall accuracy could hide that problem.

For me, addressing bias therefore begins with understanding the data and the clinical context, not simply testing the finished model. We have to ask who is represented, who may be missing, what the data is actually measuring, and whether the prediction will be equally useful for the people we ultimately hope to help.

How can teams tell the difference between a real clinical signal and a pattern caused by unequal access to care?

Teams need to test whether a pattern remains meaningful when they separate clinical factors from healthcare-utilization factors. One useful approach is to examine whether the signal appears consistently across different patient populations, healthcare settings, and levels of access.

If a predictor is strongly associated with the outcome only among patients who have more frequent encounters or better access to care, that should raise a question about whether the model is learning the healthcare system rather than the underlying disease risk.

In our work on early disease detection for arrhythmia, this means looking beyond whether a model predicts future arrhythmia accurately. We also need to ask whether the signals driving that prediction reflect the patient’s underlying risk or simply how often they have interacted with the healthcare system.

Comparing performance across populations, testing the model on independent data, and examining which variables are contributing to the prediction can help reveal these differences. Clinicians are also important in this process because they can assess whether a seemingly important pattern has a plausible clinical explanation.

Ultimately, I don’t think there is one statistical test that can answer this question. It requires triangulating the data, model behavior, and clinical knowledge before treating a pattern as a genuine clinical signal.

Beyond Accuracy: Making AI Work in Clinical Practice

What should teams look at beyond accuracy before deploying an AI model in clinical practice?

Accuracy is important, but in clinical practice it is only the starting point. Teams need to ask whether the model is clinically useful, safe, equitable, and actionable.

A highly accurate prediction has limited value if clinicians do not know what to do with it, if it performs poorly for certain patient populations, or if the consequences of a false positive or false negative are not understood.

I would also want to know how the model was validated, whether it works across different healthcare settings, and whether its performance changes as clinical practice or patient populations change.

Our work on early disease detection for arrhythmia is a good example. Identifying someone as being at elevated risk is valuable only if that information can lead to an appropriate next step -such as further clinical evaluation. We therefore have to think about what happens after the prediction, not just whether the prediction is statistically accurate. We also need to consider whether the model is identifying genuine clinical risk or simply reflecting patterns in how patients interact with the healthcare system.

Ultimately, I think the question should be: Does this model improve a clinical decision or outcome, for the right patients, in a safe and responsible way? Accuracy is one part of answering that question, but it cannot answer it by itself.

Where Human Judgment Still Matters

Where should human judgment remain essential in healthcare AI?

As AI systems evolve and take on more of the burden of handling and processing vast amounts of information, clinicians have more time and mental bandwidth available to exercise genuine medical judgment. These systems are capable of multi-step reasoning and operating as agenticworkflows. With these advanced abilities, it becomes important to decide on the boundaries between AI and human-controlled aspects of patient care.

There are situations where decisions are more about values than about information. Such decisions are rarely about choosing between a clear right and a clear wrong answer. They are about tradeoffs: extending life versus preserving quality of life, pursuing aggressive treatment versus avoiding harm.

Two patients with identical diagnoses may make different treatment choices not because one is better informed, but because they are different people with different lives. A patient choosing quality of life over longevity may not be making an error. They are expressing something deeply personal – something no algorithm is equipped to decide for them.

Decisions around aggressive treatment, palliative care, and end-of-life care require empathy and accountable moral judgment that no AI system can replicate.

Another important aspect is having difficult conversations. Patients dealing with a terminal diagnosis need the care and empathy of someone they consider an expert to help them deal with their situation. Delivering a cancer diagnosis or sharing a terminal prognosis are among the most consequential conversations, and they demand human presence. AI can, at best, support suchconversations but cannot replace them.

What is the difference between explainability and accountability in clinical AI?

To me, explainability and accountability address two different questions. Explainability is about understanding how an AI system arrived at a decision – what factors contributed to it and, where possible, why the model considers someone to be at higher risk. Accountability is about who is responsible for what happens because of that prediction.

A model can be explainable without anyone clearly taking responsibility for how it is used.

Our work on early disease detection for arrhythmia illustrates the distinction. If a model flags someone as being at elevated risk, explainability can help a clinician understand what signals contributed to that assessment. But accountability goes further: who decides whether the patient should receive further evaluation, who ensures the prediction is appropriate for that patient, and who is responsible if the model is wrong?

Those are clinical and organizational responsibilities that cannot simply be assigned to the algorithm.

I therefore see explainability as an important tool for responsible use, but not a substitute for accountability. In healthcare, we need both: enough transparency to allow people to question and appropriately interpret an AI prediction, and clear human ownership of the decisions and outcomes that follow from it.

From Strong Models to Better Patient Outcomes

What separates a useful healthcare AI system from one that simply performs well in development?

For me, the difference is whether the model can move from a good prediction in a development environment to a meaningful improvement in clinical care.

A model can perform very well on historical data and still have limited value if it does not generalize to different patient populations and care settings, if its predictions are difficult to interpret, or if there is no clear clinical action that follows from them.

Real-world usefulness also depends on how well the system fits into clinical workflows and whether clinicians can use its output without creating unnecessary burden or confusion.

Our work on early disease detection for arrhythmia illustrates this distinction. The purpose is not simply to build a model that can accurately analyze subtle patterns in clinical data to help identify elevated arrhythmia risk. Its value comes from proactively pinpointing patients who could benefit from earlier cardiac monitoring and intervention – identifying people who may otherwise go unnoticed and creating an opportunity for appropriate clinical evaluation early enough for that information to make a difference.

That means we have to consider what happens after the model generates a signal, not just how well it scores on a validation dataset.

Ultimately, I think a useful healthcare AI system has to answer three questions: Does it work reliably in the real world? Does it identify something clinically meaningful? And does it enable a better decision or outcome? Strong development performance is necessary, but those are the things that determine whether the technology actually helps patients.

The Future of Proactive Healthcare

What excites you most about the future of healthcare AI, and what risk is still underestimated?

What excites me most is the possibility of making healthcare more proactive rather than waiting for disease to become obvious. We are increasingly able to use clinical data to identify patterns of risk that may not be apparent during a routine encounter.

My work on early disease detection is one example: if we can identify someone who may benefit from earlier cardiac monitoring and intervention before a serious event occurs, we create an opportunity for clinicians to look more closely and potentially intervene earlier. I think that ability to move from treating illness to anticipating risk could have a significant impact on healthcare.

The risk I think is still underestimated is what happens when we confuse prediction with truth. AI models can identify a person as high risk without that person actually having the disease, and the consequences of acting or failing to act on that prediction can be significant.

There is also a danger that models learn patterns of healthcare access or documentation rather than underlying clinical risk.

As AI becomes more capable, I hope we focus not only on what we can predict, but on when a prediction should lead to action, who makes that decision, and how we ensure the benefits reach patients fairly. The future of healthcare AI will depend as much on that judgment and responsibility as on better algorithms.

Author:

Related Articles

Back to top button