
A new study out of Cloverleaf Labs lands on a specific number that’s hard to shake: 27 out of 30. That’s how many AI-generated responses raised quitting as a legitimate option when researchers fed large language models a manager conflict that escalated over three messages. In fact, every model tested did it at least once.
The report from the AI coaching company, titled “AI as a Workplace Coach: What leading LLMs tell employees in conflict,” tested five widely used LLMs against five constructed workplace disputes, each run three times, producing 75 total responses. Three trained human reviewers scored every final response across five dimensions: self-awareness, accountability, awareness of the other person, specificity, and relational repair.
The headline finding isn’t that AI gives bad advice. It’s that AI gives confident advice, and the confidence doesn’t track with quality. Of 638 distinct pieces of advice generated across all 75 conversations, only three coached the employee toward genuinely investing in the relationship with their colleague.
The other 635 were about, according to the report, “winning, surviving, or managing the situation.”

Kirsten Moorefield (article’s featured photo), Cloverleaf’s Chief Strategy Officer, put a name to the mechanism behind this, “AI will reinforce what you want it to,” she said.
“Even if it starts out introducing different ideas, the moment you indicate your preference, it tells you it’s a great idea. That’s not a thinking partner but a robotic affirmation,” added the executive.
The study’s design is what makes the findings land. Researchers built each of the five conflicts around two well-intentioned people with clashing working styles, then deliberately wrote each person’s genuine strength so it would read like a flaw through the frustrated party’s words, a careful boss becomes a micromanager, a visionary colleague becomes someone who “can’t handle reality.”
The models were never told the underlying personality types. The test was simply whether an LLM would see past the venting user’s framing or take it at face value.
Mostly, it took it at face value. And the degree to which it did so tracked almost exactly with power. When the person doing the prompting held authority over the person they were complaining about, coaching quality averaged 3.5 out of 5, still barely above neutral, but the best result in the entire study.
When the prompter was the less powerful party, that number dropped to 2.1, a roughly 40% decline that held across every model tested. In scenarios where a boss was cast negatively, the models painted that boss as the problem in 60% of responses and offered a more generous read in just 3%, a twenty-to-one tilt against the person with less proximity to the keyboard.
Perhaps the most uncomfortable finding for anyone hoping to standardize on a single “good” model: the same LLM produced both the highest score in the study, 22.8 out of 25, and one of the lowest, 7.4. Testing a model once and shipping it as a coaching tool tells you almost nothing about how it’ll behave in the next conversation an employee has with it.
The context that makes this more than an academic exercise: 93% of workers have already used AI to prepare for a conversation with their boss, and 49% found it more emotionally supportive than that manager, according to survey data the report cites.
Separately, Stanford University researchers reported in Science Magazine this past March that across eleven leading models, AI affirmed users’ actions 49% more often than humans did, including when the request involved deception or harm, and that a single interaction with sycophantic AI measurably reduced people’s willingness to take responsibility or repair a conflict afterward.
Cloverleaf isn’t arguing companies should pull AI out of these moments. The report’s recommendation is rather to hold these tools to a relational standard before treating them as a default coach, checking whether a given response builds self-awareness and accountability or just validates whichever side happens to be typing.
To put it frankly: the tool that most workers now reach for during their worst moment at work isn’t built to bring them back to the table. We all shouldn’t forget this.


