AI & Technology

Five Real Sources Out of Forty-Five: What the Big Four Just Taught Us About AI Research

By Rob Smith

Two of the world’s largest advisory firms withdrew reports after AI invented the evidence. The model behaved exactly as designed. The process around it did not.

In October 2025, KPMG published a report called “Total Experience: Redefining Excellence in the Age of Agentic AI,” and in June 2026 it took that report down. The reason is almost too neat: research arguing that AI is ready for serious work had, in significant part, been written by AI that made things up.

I have spent most of my career producing research that buyers act on, so I want to be careful about what this story is. It is not evidence that AI cannot be trusted with research. It is evidence of something narrower and far more useful: what happens when nobody checks.

What actually happened

UBS, the UK’s National Health Service, Swiss Federal Railways and Transport for London all told the Financial Times that the report’s claims about their AI use were wrong or misleading. One example stands out: the report described an Emirates mobile chatbot called Sara that rebooks flights, when Sara is a robot assistant introduced in 2023 that cannot change a booking.

Then the detection firm GPTZero went through the citations and found that of the report’s 45 sources, five pointed to real, intact material, while 40 of the titles were fabricated outright. GPTZero gave the behavior a name: vibe citing, where a model stitches fragments of genuine sources together and invents the rest.

Weeks later EY Canada withdrew a 44-page cybersecurity report, “Points of Attack: Uncovering Cyber Threats and Fraud in Loyalty Systems,” after the same firm found that 16 of its 27 cited sources were fabricated, misattributed or pointed at dead links.

Two of the Big Four. The same failure. Weeks apart. That is not bad luck 

The model did nothing unusual

Here is the part that keeps getting misreported. A language model does not retrieve a citation and judge whether to trust it; it predicts what a citation for that claim would probably look like and then writes it, complete with a plausible title, a plausible author and a plausible year. The output is shaped like a source. It is not one.

This gets worse in fast-moving subjects, where the real material is recent and thinly represented in whatever the model learned from, and agentic AI in 2025 is exactly that kind of subject. So is fraud in loyalty systems. Both reports chose topics where the failure mode is at its most aggressive.

The wrong answer also looks precisely like the right one, because there is no formatting difference between a verified citation and an invented one. No hedge, no flag, no tell. That is what makes it dangerous inside a document that hundreds of executives will read and nobody will spot-check.

Why it landed hardest on advisory firms

Any organization could have shipped this, but it matters more when a trust business does.

What these firms sell is judgment, and a client pays for the conclusion because of who stood behind it rather than because of how it reads. When the verification step disappears, the client is paying for formatting.

I would rather the industry treat this as a design problem than a discipline problem. Nobody at KPMG or EY set out to publish invented sources; the workflow simply had no point at which a person with domain knowledge was required to confirm that the evidence existed before it shipped. Add enough speed to a process with a gap like that and the outcome stops being a surprise.

The fix is sequencing, not abstinence

I use AI in research every day and I am not going to stop, and neither is anyone reading this. Telling people to avoid the tools is advice nobody follows, and it is poor advice anyway.

The fix is order of operations. AI is very good at collecting, drafting and summarizing at a speed no human matches, and unreliable at establishing whether a thing is true. So put the person where the truth claim gets made, not where the prose gets polished.

Three questions handle most of it.

Who verified this? Not who reviewed the writing, but who opened the sources.

Does the cited source exist, and does it say what we claim it says? Both halves matter, because a real paper attached to a claim it never made is a different error with exactly the same consequence.

If a named organization appears in our evidence, would they recognize the description? UBS and Transport for London did not, and they said so in public. That check costs one email.

None of this is exotic. It is the fact-checking any serious publisher did routinely before 2023, and what changed is that the volume of confident, well-formatted, unverified text rose by an order of magnitude while most verification processes did not move at all.

What buyers should take from it

If you commission research, or pay for it, you now have a fair question for any provider: at what point in your process does a person confirm the evidence, and who is that person?

Vagueness in the answer is the answer. A firm doing this properly can name the step and the role in a single sentence, and a firm that cannot is describing hope.

Both reports came down. The lesson is not that AI wrote them. The lesson is that nothing stood between the model and the client.

Related Articles

Back to top button