AI & Technology

One in Ten AI Interactions Fails. Your Test Suite Cannot Tell You Which Ones.

By Darin Brown, Chief Product and Technology Officer, Testlio

Most quality conversations still assume failure looks like breakage. Something crashes, a call returns a 500, a test goes red. That model has held up for thirty years. 

It does not hold up for AI. 

Over the past twelve months, we ran a study on AI failures. Our trained community testers ran thousands of structured exploratory prompts against multiple product-specific AI assistants, meaning deployed customer-facing bots rather than frontier models, and scored every result by hand. Ten percent of interactions failed. Almost none of those failures would have been caught by automated evaluation, because we were scoring behavior rather than whether a response came back. 

Ten percent is the number I would put in front of a board. 

The failures that matter are not the loud ones 

Across all our client engagements over the same period, roughly 60% of the issues we found were medium or high severity. In the AI study, that figure was 77%. These are different populations and the comparison is directional rather than controlled, but the direction matches what we see in the field. When AI fails, it fails in ways that matter more often than conventional software does. 

Within the high-severity slice, the distribution is the part worth studying. 

Forty-six percent traced to safety guardrails and fallback handling. Inconsistent refusals, over-permissive answers to edge-case phrasing, contradictory logic when a conversation went somewhere unscripted. 

Thirty-nine percent were output accuracy and intent resolution. The assistant answered a plausible question that was not the one asked. 

Thirteen percent were hallucinations and misinformation. Invented policies, fabricated figures, confidently wrong procedural guidance. 

Guardrails are the largest category. That should bother anyone who licensed a model with safety claims attached and treated the matter as settled. 

Guardrails do not transfer 

Here is the argument I would make to any executive shipping AI into a customer path. 

A guardrail is not a feature you inherit from your model provider. It is a judgment about your business. What your assistant should refuse, what it should escalate, what it should never assert about a refund or a diagnosis or an interest rate, all of that is specific to your policy, your regulator, and your customers. No foundation model provider knows any of it. 

The risk does not transfer with the purchase. You can buy a model with strong general safety behavior and still ship a guardrail failure on day one, because the guardrail that matters is the one nobody outside your company could have written. 

Our data says that is where nearly half of serious AI failures live. 

Your users will not tell you 

The second thing making this hard is that the failure is invisible to the people experiencing it. 

Kantar found that 30% of global consumers say they often or always get incorrect or misleading answers from AI, and 23% say they cannot tell when AI is wrong. Together those describe a system with no feedback loop. Users are absorbing bad output at scale and a quarter of them cannot flag it because they cannot detect it. 

Meanwhile 7 in 10 consumers say they would take their business elsewhere after a single bad AI experience. They will not file a ticket. They will leave. 

Agentic systems make this worse. An agent can complete its task and still violate a policy, expose data it should not have touched, or make a decision that harms the customer, while every log looks clean. The functional trace shows success. 

What to do about it 

Five things, in rough order of what I have seen actually move the number. 

Write the guardrail spec before you write the prompt. Not a list of banned topics. A document stating what the system must refuse, what it must escalate to a human, and what it must never assert without a citation. If you cannot write that document, you are not ready to deploy. 

Test adversarially, with humans scoring. Automated evaluation catches drift and statistical anomaly. It cannot judge whether a response was subtly misleading, or whether a guardrail held under a hostile rephrase. That takes human judgment against a rubric, calibrated across reviewers. Without the rubric and the calibration you are collecting opinions, not evidence. 

Assign decision rights explicitly. Engineering owns where guardrails sit architecturally. Product owns which actions the system takes autonomously, which need sign-off, and which it does not take at all. Design owns whether the user can tell what the system is doing and has recourse when it is wrong. If those are not named people, nobody owns it. 

Close the access control gap. IBM reported that 13% of organizations had a breach of an AI model or application last year, and 97% of those lacked proper AI access controls. That is a governance failure, not a testing failure, and no amount of evaluation compensates for it. 

Instrument for silent failure. Conventional monitoring tells you when something broke. You need to know when nothing broke and the outcome was still wrong. That means sampling and reviewing successful interactions, not only failed ones. 

The honest version 

I will state the limits of our own data. The AI study focused on chatbot experiences. It reflects our testing population, not a market sample. Anyone publishing research about their own customers, including us, deserves that scrutiny. 

But 10% is not a rounding error, and 46% concentrated in guardrails is not noise. 

Cost is not the constraint here. Splunk puts unplanned downtime at $15,000 per minute, with aggregate Global 2000 losses reaching $600 billion annually, up 50% over two years. Forty percent of companies already estimate poor software quality costs them more than $1 million a year. The money is being spent either way. 

The question is whether you find the failures or your customers do. With AI, your customers may not know they have. 

Findings referenced are from Testlio’s 2026 Software Quality Report, drawing on Testlio client engagements over twelve months and a separate AI testing study. 

Related Articles

Back to top button