
More business decisions now start with an AI answer. In a March 2026 survey of 1,076 B2B software buyers and decision-makers, G2 found that 51% start their research with an AI chatbot more often than with Google, up from 29% in its 2025 buyer behavior report.
Most of the concern about that shift is hallucination: the model making something up. We wanted to test a quieter problem, one that is easy to miss because the answer sounds right. When an AI answer hands you a fact, can you see where it came from, and does that source actually say it?
So on September 29, 2026, we ran a test at Sourcelift. We asked ChatGPT, Claude, Gemini, Perplexity and Google’s AI Overviews the same 16 questions, then checked all 80 answers against the primary source. The definitions were right. The sourcing wasn’t.
How We Ran the Test
The questions covered eight terms from AI search, among them GEO, llms.txt, query fan-out and E-E-A-T, each asked once in English and once in Spanish. We used each engine’s consumer app, logged in, with its default model and settings. For Google, we searched google.com and read the AI Overview on the results page.
We read every answer in full and checked the definition against the primary source: the originator’s own documentation or paper, such as Google Search Central for AI Overviews or the original arXiv paper for GEO. We also flagged any claim that was unsourced, outdated or contradicted by that source.
It is a small sample, with one run per question on a single day. AI answers vary from run to run, so we read the results as signals rather than as a ranking of engines. The prompts, result tables and notes are published in full in our GEO glossary.
What We Found
The definitions were right. All 80 answers defined the term correctly. On the basics, today’s engines are reliable, which is exactly why the problems below are easy to miss.
Most answers skipped the primary source. Leaving out Google’s AI Overviews, only 21 of the 64 answers from ChatGPT, Claude, Gemini and Perplexity linked a primary source. The AI Overviews cited a Reddit thread, a Facebook post and an Instagram post, and the Spanish one about llms.txt did not link Google’s own guide on the subject.
Precise numbers came with nothing behind them. Asked about query fan-out, the technique AI search uses to split one question into several searches, Google’s AI Overview said it typically runs 5 to 12 searches. Gemini, asked in Spanish, said between 5 and more than 20. Google’s documentation describes the technique but gives no number.
A bigger claim got the same treatment. Perplexity, answering in Spanish, repeated an SEO blog’s claim that close to half of all searches already happen inside AI answers, with no data behind it. The closest independent measure we found is Pew Research Center’s: in browsing data from 900 US adults in March 2025, 18% of Google searches produced an AI summary.
Facts went stale, even about the engine itself. Google’s AI Overview said AI Overviews are available in more than 120 countries. Google had announced more than 200 countries and territories on May 20, 2025.
Advice contradicted the platform it described. In both languages, Gemini recommended FAQPage and HowTo markup and answer blocks of 40 to 60 words for answer engine optimization. Google’s guide says there’s “no special schema.org markup you need to add”, and FAQ rich results stopped appearing in Google Search on May 7, 2026.
Origins went missing. Only Claude credited the 2023 paper that introduced GEO. None of the 10 answers about retrieval-augmented generation credited the 2020 paper that named it.
Language changed the answer. In English, 14 of 32 answers linked a primary source; in Spanish, 7 of 32. In Spanish, two engines also called Google’s AI Overviews by names Google doesn’t use, rather than its official Spanish name.
Different accounts, different answers. Twelve of 16 ChatGPT answers were tailored to the account, drawing on its business, location or earlier chats, even though we used temporary chats. Perplexity answered 4 of 8 English questions in Spanish, the account’s interface language.
Why “Mostly Right” Is Still a Risk
The engines did their core job well. What they got wrong sat in the details that make an answer sound authoritative: a number, a date, a country count, a best practice. Those are the details that end up pasted into strategy decks and board papers.
An unsourced claim isn’t necessarily wrong, but it is unverifiable. Nobody downstream can trace where it came from, or tell when it stopped being true.
Our sample is small, but larger studies point the same way. In a study coordinated by the European Broadcasting Union and led by the BBC, journalists from 22 public service media organizations reviewed more than 3,000 answers to news questions from ChatGPT, Copilot, Gemini and Perplexity, and 31% had serious sourcing problems: missing, misleading or incorrect attributions. A link is not proof either. In a 2023 audit of four generative search engines, only 74.5% of citations supported the sentence they were attached to.
Meanwhile, few people check. In Pew’s data, users clicked a link inside Google’s AI summary in just 1% of visits to pages that showed one. For most readers, the answer is the product and the source is a footnote.
Five Habits for Teams That Decide With AI
- Ask for the primary source behind any number that goes into a document. That means the paper, standard, filing or official documentation the claim comes from, not a blog that repeats it. OpenAI’s own help page for ChatGPT search puts it plainly: “Open a cited source to check that it supports the answer.”
- Treat precision without a source as a red flag. “5 to 12 searches” sounds like research. If the engine can’t show where a number comes from, treat it as a guess until someone confirms it.
- Check the date before you trust the fact. Platform facts change quickly. In our test, one engine repeated a country count that was 16 months out of date, and another recommended markup for a Google feature retired months earlier.
- Ask in the language you’ll act in. If a decision spans markets, run the question in each language and compare. In our test, the Spanish answers linked primary sources half as often as the English ones.
- Use clean sessions for anything you benchmark. Logged-in answers are personalized, and they shift between runs: in a SparkToro study of 2,961 runs, there was less than a 1 in 100 chance that ChatGPT or Google’s AI would return the same list of brands twice. If you compare vendors or track how AI describes your company, repeat each prompt in fresh sessions.
The Other Side of the Gap: Be the Source
For brands, the same gap is an opening. When an engine can’t find a clear primary source, it fills the space with whatever it can retrieve, which in our test included forum threads and social posts. If your company knows a subject best, the most useful thing it can do is make sure the primary source exists, and that AI systems can find it.
We see what happens when that step is skipped. In a September 2026 diagnosis, a small local business was named in 9 of 25 AI answers, yet none of those answers linked to its website; they linked to third-party listing sites instead. One reason was mundane: a noindex tag on every page had kept the site out of search indexes.
Substance matters more than tricks here. A 2025 benchmark found most content-rewriting tactics for conversational search “largely ineffective”, and Google’s guide favors what it calls non-commodity content, which “provides unique expert or experienced takes that go beyond common knowledge and the ordinary.” Clear definitions, dated facts, original data and named authors, on pages AI crawlers can read, are what make a page worth citing. That is the real work of generative engine optimization (GEO): becoming the primary source on your own subject.
Correct Is Not the Same as Verifiable
AI answers are becoming the first draft of many business decisions, and on the basics they are good. The risk sits in the details: the number with no source, the fact that expired last year, the advice the platform itself has dropped.
So keep using AI, and hold its answers to the standard you would set for a new analyst: show your sources. And if your expertise is the answer people are looking for, make sure AI has a primary source to find.
About the Author
Tony G. is the founder of Sourcelift. Sourcelift is a productized GEO and AEO agency that gets ecommerce, legal tech, health, fintech and SaaS brands cited inside AI answers. The full engine test, with all 16 prompts and both result tables, is in Sourcelift’s GEO glossary, and brands can see how AI engines describe them with a free AI visibility diagnosis.

