
Customer service agents at Alibaba who were given a generative AI assistant answered faster and earned higher customer ratings. Customers still came back to retry their issue as often as before, according to a large-scale field experiment in the company’s after-sales chat support, first posted in February 2026.
The assistant drafted an issue diagnosis and a proposed solution in the opening stage of a chat, and agents could adopt or ignore it. Lower-performing agents gained the most. Among top performers, using the assistant was tied to slower responses and more customers retrying right away, a pattern the authors say is consistent with workflow disruption.
The study did not examine how current the assistant’s information was, a question other research this year has taken up.
Outdated Information Can Override a Correct Answer
A single outdated document overturned 30% to 37% of the answers two AI models had already gotten right, even when the models were not told to trust it, according to a study from the University of Southern Denmark posted in September 2026. The researchers tested retrieval-augmented generation, or RAG, in which a model looks up reference documents before it answers. Many customer service AI tools use the same basic approach to answer from a company’s knowledge base.
The team assembled 317 cases where the correct answer had changed, across medicine, law, software and platform policy. When the models were instructed to follow the outdated document, the flip rates rose to 66% and 75%. Across the four open-weight models and four domains tested, they ranged from 17% to 91%. Given the current version of the same document instead, the models followed it in 97% to 100% of trials.
In the tests, the models followed current documents almost every time, while outdated ones still overrode a substantial share of answers the models had gotten right. The authors recommend that retrieval systems record when guidance stops applying, such as effective dates or supersession relations, because a publication date alone did little to stop models from using outdated evidence.
The Best Contact Center Knowledge May Never Reach a Manual
In a contact center, some of the most current knowledge sits with the agents who handle the calls. “Agents improve when training, assist, quality and coaching operate as one connected system, because that is the only way the operation learns from its own best work,” said Erin Walker, Global Vice President of CX AI, Business & Delivery at TELUS Digital in TELUS Digital’s 2026 buyer’s guide to contact center AI partners.
The lower-performing agents who stand to gain from that knowledge are the same group the 2026 field experiment found AI helps most. A 2025 study in The Quarterly Journal of Economics by economists at Stanford and MIT found something similar. Following more than 5,000 customer support agents as a generative AI assistant was introduced, the authors found that newer and lower-skilled agents saw the largest productivity gains, while the most experienced saw little change. They also found suggestive evidence that the tool spread the tacit knowledge of more able workers to everyone else.
Keeping Contact Center AI Current After Launch
In an October 2026 release on AI in customer service, TELUS Digital said a frequent cause of a confident but wrong AI answer is fragmented source material, where a tool draws from “multiple, sometimes conflicting internal documents” and produces something plausible instead of flagging that it is unsure. The release recommends governing model access and content centrally, so answers draw from current and consistent sources, and running responses through automated accuracy and safety checks before a customer sees them.
“AI should take the parts of a customer interaction that slow an agent down, so the agent can spend their judgment where it counts. Get that division right for each use case and you get a better experience for the customer and a better job for the agent at the same time,” Walker said in the release.
Speed Alone Can Hide Whether AI Is Working
The release also cautions against judging customer service AI on a single efficiency number. Average handle time can rise after AI absorbs simpler requests, because the work left for agents is harder. Managers should look instead at first contact resolution, customer satisfaction and the quality of the interactions that reach agents.
The Alibaba researchers recommend that companies evaluate performance carefully and tailor rollout to agent skill. Their study measured speed, ratings and repeat contacts. It did not measure how current the assistant’s knowledge base was.
