
Before joining Microsoft, I spent time at FSLogix, watched Microsoft acquire it, then spent more years inside Microsoft building Azure Virtual Desktop and Windows 365. If there’s one pattern I’ve seen repeat itself more than any other, it’s this: organizations roll out a new capability on top of an environment they never actually finish fully optimizing.
I watched it happen with VDI. I watched it happen again with the move to Azure. And I’m watching it happen right now with AI.
The pattern is always the same. A workload that was tuned, sized, and budgeted for yesterday’s usage gets asked to absorb something fundamentally different in shape. And the cracks that were always there, just quiet, suddenly aren’t quiet anymore. AI workloads don’t ease into an environment. They spike hard, scale unpredictably, and demand elasticity that an under-optimized environment was never built to deliver. The infrastructure doesn’t fail loudly. It just quietly produces worse results, higher bills, and a pilot that gets killed six months later for reasons nobody can quite articulate.
After years of working at the intersection of cloud infrastructure and enterprise computing, I’ve stopped believing AI initiatives fail because of bad strategy. They fail because the foundation was never ready to begin with. Three fault lines show up again and again: cloud environments that haven’t been optimized, data that can’t be trusted, and cost structures that spiral the moment pilots go live.
Gartner has warned that many GenAI projects are being abandoned after proof of concept because of poor data quality, inadequate risk controls, escalating costs and unclear business value. These are the same operational issues that tend to get overlooked when organizations move too quickly from experimentation to production.
You’re Probably Paying for More Than You’re Using
Before any organization adds AI to its environment, it should ask a simpler question: are we actually using what we have?
Most aren’t. The lift-and-shift migration era left a lot of organizations running legacy patterns inside modern cloud environments. The compute is there, the tooling exists, but the configurations, like instance sizing, auto-scaling policies, and reserved capacity, haven’t kept pace. The result is an environment that’s simultaneously over-provisioned in some areas and too rigid in others.
The work here isn’t glamorous. It means auditing whether virtual machines are sized for actual usage, not historical estimates. It means revisiting auto-scaling rules that were set during migration and never touched again. None of this is cutting-edge, but done consistently, it creates the headroom AI workloads need, and establishes the efficiency baseline that makes costs predictable when demand scales.
The Data Problem Is Worse Than You Think
If infrastructure efficiency is the floor, data quality is the foundation. And in most organizations, the foundation has cracks.
Gartner’s research reinforces that point: 63% of organizations either do not have or are unsure if they have the right data management practices for AI, and Gartner predicts that through 2026, organizations will abandon 60% of AI projects that are not supported by AI-ready data.
IBM has flagged data quality and readiness as one of the largest barriers to enterprise AI adoption, noting that fragmented data, inconsistent formats and weak governance can directly undermine AI performance and reliability.
AI systems — whether they’re powering intelligent automation, endpoint management, or user experience optimization — draw conclusions from the data they’re given. Feed them stale, incomplete, or inconsistent data, and the outputs are either wrong or worthless. Often both.
Three dimensions determine whether data is actually ready for AI work. The first is accuracy: are records current, consistent, and free of the kind of accumulated errors that build up when no one’s actively governing a dataset? Device profiles that haven’t been refreshed in six months. User attributes that don’t reflect org changes. Policy assignments that are technically active but practically meaningless.
The second is completeness. An anomaly detection system can’t establish a normal baseline if your telemetry has gaps. A recommendation engine for application delivery can’t surface useful signals if most of your device data is missing. Before deploying AI features, it’s worth mapping the data those features require against what you genuinely have.
The third, and the one organizations consistently underestimate, is timeliness. Batch pipelines were adequate for a past generation of reporting and analytics. They’re not adequate for AI systems that need to act on what’s happening now. If the intelligence arriving at your AI layer is twelve hours old, its recommendations are twelve hours behind reality. In fast-moving environments, that gap compounds quickly.
Data infrastructure work — cleaning pipelines, instrumenting telemetry, establishing governance — rarely makes it into board presentations. It’s not the kind of investment that generates excitement. But it’s the difference between AI features that earn lasting adoption and AI features that get quietly retired because no one trusted the outputs.
Cost Is an Architectural Decision, Not a Finance Problem
The third failure mode is financial, and it tends to arrive as a surprise.
AI workloads are expensive and difficult to predict. A pilot that runs efficiently at limited scale can generate dramatically different cost curves in production, especially in environments like virtual desktops, where AI-powered features like session intelligence, predictive scaling, and automated remediation operate simultaneously across thousands of endpoints.
Organizations that navigate this well treat cost governance as something they design before AI workloads arrive, not something they respond to afterward. That means establishing per-user, per-session, and per-workload cost baselines during the optimization phase, before AI is layered on, so that any anomalies AI introduces are immediately visible against a known benchmark.
It also means choosing infrastructure management tools that surface cost signals in real time. Monthly cloud bills are too slow. By the time an overrun shows up there, weeks of unnecessary spend have already happened.
What Comes Next Is Harder, Not Easier
Here’s the part most of these conversations stop short of, and it’s the part I think matters most.
Everything above assumes the thing consuming your infrastructure is a person, or at least a workload acting on a person’s behalf. That assumption is already breaking down. As agentic AI matures from pilot programs into production deployments, infrastructure readiness has to expand to cover AI agents as first-class citizens of the environment.
Agents need their own compute. Their own identity and access boundaries. Their own governance models, running alongside human users rather than borrowing capacity meant for them. Most organizations haven’t even finished the human-side optimization work described above. Very few have started thinking about what happens when the thing consuming compute, generating cost, and touching data isn’t a person at a desk, but a system acting autonomously, around the clock, at a scale no IT team has had to plan for before.
That makes the foundation work more urgent, not less. The gap between AI ambition and AI execution isn’t structural, it’s operational. It reflects choices that can be revisited: cloud environments that can be optimized, data pipelines that can be modernized, cost frameworks that can be built with governance in mind from the start.
The organizations I’m most optimistic about aren’t necessarily the most technically sophisticated. They’re the ones willing to be honest about where they actually stand before committing to where they want to go.



