
A recent Anthropic analysis of more than 81,000 Claude conversations turned up a surprising number. Personal transformation ranks among the top reasons people turn to AI, but only 6% of users say it gives them emotional support. We’ve built the smartest assistants in human history, but we’re failing at one of the things people actually want them for.
I’ve spent over a decade building one of the largest psychology organizations in the world. I see the human cost of this failure every day. When you hand someone an easy answer, the thing most AI platforms strive to do, you rob them of the chance to grow. That 6% figure confirmed the problem we’ve been working to resolve throughout the history of psychology.
The industry has spent the last several years racing to make AI smarter. Every few months brings a new benchmark score and a new claim about reasoning ability or coding proficiency. While valuable, those gains miss the broader humanistic picture.
Answers Aren’t the Same as Judgment
Today’s models are extraordinary at producing answers. Ask a question and you’ll get something fast, confident, and comprehensive. For a huge range of problems, that’s exactly what you need. If your server is throwing errors, you want the fix. If you’re trying to understand a legal filing requirement, you want the rule, not a reflective question about your feelings toward paperwork.
But a large share of what people bring to AI isn’t that kind of problem. Should I take this job? Should I stay in this relationship? How do I even begin to process what just happened in my family? These questions don’t have a correct answer sitting in a database somewhere. The relevant information lives inside the person asking, built from years of context that a model can’t access or replicate.
When Task Aligned AI treats every question the same way and defaults to comprehensive resolution regardless of what’s actually being asked, it starts to create a strange dependency loop. The person gets worse at trusting themselves because they’re outsourcing decision-making. Over time, users can lose their own judgment
This is not breaking news in the field of psychology. Albert Bandura’s research on self-efficacy found that people build confidence by acting and seeing the results of their own effort, not by being handed conclusions. When someone works through a hard decision and is ready to live with whatever’s the outcome, they grow far more from the experience than when an outside authority resolves it for them.
We’re Measuring the Wrong Thing
Part of the reason this problem persists is that we’re not measuring for it. Ask any major AI lab what “alignment” means and you’ll hear about safety. Does the model refuse to help someone build a weapon? Does it avoid encouraging self-harm? These are essential guardrails, and I have no issue with the labs building them.
But they only describe the baseline. They tell us what a model won’t do in the worst case. They tell us almost nothing about what it does in the best case, when someone brings it a hard, human problem and needs something more than a fast, polished answer.
You can see this clearly in how these AI labs talk about their newest models. In the media and on their web sites, all of the data is around the model’s effort to solve a problem or retrieve an objective fact or do a task. None of them tell us whether an interaction left a person more capable of thinking clearly on their own. That consideration for human flourishing is not just of prime importance, but also a critical safety issue in ensuring that humanity remains in the driver’s seat as it relates to our trajectory.
Meanwhile the largest AI companies are showing us their priorities with each new model release and the benchmarks they trot out, like improvements in agentic coding and novel problem solving. There’s no mention of things like adaptive scaffolding, deep listening, or modeling with agency.
We built a framework called the Sovereign Human Benchmark to try to address that. It evaluates whether a conversation leaves someone more connected to their own thinking or more impressed by the AI’s. Each of the 12 criteria are grounded in decades of psychological research concerned with methodologies by which we can accelerate human potential. This critical conversation is currently missing from the discourse, and the existential risks of not addressing the degree to which AI is aligned to humanity could prove catastrophic.
What a Human-Centered Approach Could Look Like
We don’t have to abandon the progress AI has made. However, expanding the understanding of the various roles AI can fill is a communication failure on the part of the industry. Specifically: Task Aligned AI and Human Aligned AI are not the same thing.
Human Aligned AI knows the difference between an objective question and a subjective one, and responds accordingly. Ask about a medication interaction, and it should give you a direct, clear answer. Ask whether to leave a stable job for an uncertain opportunity, and it should slow down, ask what actually matters to you, and help you find your own answer instead of supplying one. Knowing when to hold back an answer is as much a skill as knowing when to give one, and a more human-aligned model would value restraint.
It would also measure success differently. What happens during the conversation isn’t as important as what the person gained from it. Did they leave more capable of making the next hard call themselves? Did they take a step forward in trusting their own judgement?
The Opportunity in Front of Us
That 6% figure from the Anthropic study gives AI innovators an opportunity. The industry has proven, repeatedly, that it can build systems of remarkable capability. The next test is whether we can build systems that make people more capable too.
This is very possibly the beginning of the golden age of humanity. As we reach further into the cosmos do we want to be in the sidecar, or do we want humans driving the frontier? I don’t want humanity to just be along for the ride.
We already know how to measure whether AI can outthink us. It’s time we got serious about measuring whether it can help us think better for ourselves.
