Enterprise AI

When AI Stops Being an Experiment, It Starts Being Infrastructure

By Darren Kimura, CEO and President of AISquared

In June 2026, Microsoft did something that would have seemed unlikely a year earlier: it reportedly began adding Amazon Web Services capacity to support GitHub, its developer platform, after AI-driven coding growth strained GitHub’s infrastructure.¹ 

What triggered this was not ordinary software growth. AI-assisted coding, and more specifically GitHub’s commit volume, was reportedly on track to hit roughly 14 billion commits in 2026, up from about 1 billion commits the year before.² 

At the time, much of the coverage treated this as irony. But the irony is the signal. Microsoft using Amazon capacity for GitHub shows that AI demand is stressing even the largest technology platforms. When a hyperscaler needs another hyperscaler, the lesson for enterprises is clear: the AI capacity problem is real, and it could reach your organization faster than you think. 

The Conversation Is Moving From ROI to Reliability 

Inside many large companies, enterprise AI has already moved through two phases. The first was experimentation: pilots, copilots, and teams hunting for use cases. The second was ROI, when boards wanted measurable savings and CIOs wanted proof of productivity. 

Now comes the third phase: reliability. 

Reliability does not replace ROI. It changes how ROI must be measured. 

A coding assistant may help developers ship faster. A support team may use AI to summarize calls and draft replies. A finance team may use it to flag anomalies in contracts. The gains can be real. 

But the moment people build work around those tools, uptime becomes part of the business case. So does latency. So does output quality. 

Then comes the harder question: what happens when the tool is wrong? 

AI creates value when it works. It becomes a point of failure when it does not. 

AI Failure Will Not Look Like an Outage 

A traditional failure is loud. The network drops. The app will not load. Everyone knows inside a minute. 

AI fails differently. That is what makes it dangerous. 

Picture a support team whose AI keeps answering. It just stopped pulling the last two weeks of product updates, so it confidently quotes the old return policy to every customer who asks. No alert fires. The dashboards stay green. The system reports itself healthy. 

The first real signal arrives three days later as a spike in complaints and a chargeback report nobody can explain. 

The system was available the entire time. It was also wrong the entire time. 

That gap is why AI reliability cannot mean the same thing as software uptime. The question is not whether the system is online. It is whether the business can still trust the output inside the workflow where people actually use it. 

That pulls in availability, speed, accuracy, data freshness, integration health, audit trails, and human review. A system can pass every uptime check and still walk the company into a bad decision. 

AI Downtime Is Becoming Business Downtime 

Early on, an AI outage cost you an inconvenience. The summarizer broke, so someone wrote the summary by hand. 

That trade disappears once AI sits inside production. 

When engineering leans on AI coding assistants, an outage drains velocity. When support leans on AI triage, it cracks the service level. When a regulated workflow leans on AI, downtime can become a compliance problem. 

The danger is not that every use case matters today. It is that use cases become mission-critical quietly, one workflow at a time, without anyone classifying them as infrastructure. 

Business Insider reported that GitHub had already suffered dozens of major outages in 2026 as AI-driven coding demand surged.¹ Reliability gaps usually show up first as missed customer commitments. 

Business Continuity Plans Need an AI Chapter 

Mature companies already rehearse for cloud outages, cyber incidents, and vendor failures. AI belongs in that same binder. 

The first move is visibility. Where is AI used? Which workflows depend on it? Which models and vendors sit in the path? Who owns the result? 

Most companies cannot answer those four questions. 

The second move is fallback planning. When an AI system fails, does work halt, reroute to another model, or drop back to a human? Has anyone tested that path, or does the team improvise during the outage? 

Answering that before the failure costs a meeting. Answering it during the failure costs a weekend and a customer. 

Resilience requires more than a second cloud provider. It needs observability, enforcement, and a named owner. 

The Enterprise AI Stack Is Fragmenting Before Anyone Controls It 

Most enterprises already have plenty of capability. Copilots. Cloud AI services. Teams wiring up agents and retrieval systems. All of it is available today. 

The problem is not a shortage of tools. The problem is that the tools are piling up faster than anything that coordinates them. 

One team runs one model for support. Finance runs a different one. A product group ships an agent. A data team connects a vector database. Each system may work on its own. Almost none of it shares a record of what runs where. 

Shadow AI is not just the intern using ChatGPT to polish an email. That version is easy to picture and easy to manage. 

The bigger exposure shows up when disconnected workflows start shaping customer replies, code releases, and risk decisions with no audit trail behind them. 

Governance arrives late. By the time leadership sees the whole map, the company may already run on systems it cannot watch. 

Governance Describes Intent. Control Determines Reality. 

The standard response is an AI governance committee, an acceptable-use policy, and an approval process. 

Those are necessary. They are not enough. 

Governance says what should happen. Control decides what actually happens once code is in production. 

NIST’s AI Risk Management Framework is voluntary, outcome-focused, and non-prescriptive.³ It is useful, but it does not run inside production systems. It gives leaders a way to organize risk. It does not, by itself, enforce policy, produce runtime logs, or stop a failing AI workflow. 

NIST’s guidance is clear on one point: AI risk should be folded into enterprise risk management, alongside cybersecurity, privacy, and other business risks.³ 

Policy needs something underneath it that can enforce. 

A Control Layer Is Becoming Its Own Category 

The same pattern keeps surfacing in analyst notes, enterprise architecture discussions, and standards work. Companies need a layer that sits above individual models and reaches into the applications where work happens. 

In practice, that means four things. 

Visibility into what data went in, which model ran, and where the output went. Enforcement that lives in production instead of in a slide deck. Records that a risk team or regulator can trust. Cost controls that tie consumption to spend before the invoice surprises someone. 

The industry has seen this before. 

Networking scaled, and observability and management tools followed. SaaS exploded, and identity and access governance followed. Data platforms grew, and lineage and cataloging followed. 

AI is walking the same path on a shorter clock. 

The Cloud Security Alliance has published agentic AI governance work, including AAGATE, a reference architecture that translates NIST AI RMF concepts into a Kubernetes-native runtime model for continuous governance.⁵ 

A Second Cloud Is Not a Plan 

The GitHub example may suggest an easy fix: add another cloud provider. 

That may help. It is not enough. 

AI resilience is not mostly about where compute runs. A single AI workflow can lean on a model provider, a vector database, an internal data source, an identity system, an app connector, a policy layer, and a human approval step. 

Any one of them can break while every cloud stays up. 

The right question is not whether you have a backup cloud. It is whether the business keeps running when one link in that chain fails. 

Reliability and Control Are the Same Problem 

Reliability and governance often look like two initiatives. In AI, they are one problem. 

You cannot make AI dependable without knowing where it runs. You cannot govern what you cannot see. 

The next phase of AI will not go to the company with the most pilots. It will go to the one that can make AI dependable inside workflows that customers, employees, and regulators actually touch. 

Companies should not slow adoption. They should prepare for AI to matter. 

Once AI becomes infrastructure, resilience stops being an IT line item and becomes part of the return. 

References 

  • Ashley Stewart, “GitHub’s AI Surge Pushes Microsoft Into Amazon’s Arms,” Business Insider, June 2026. 
  • Craig Hale, “Microsoft forced to turn to AWS to boost GitHub cloud capacity following AI demand surge,” TechRadar, June 2026. 
  • National Institute of Standards and Technology, “Artificial Intelligence Risk Management Framework (AI RMF 1.0),” NIST AI 100-1, January 2023. 
  • National Institute of Standards and Technology, “AI Risk Management Framework,” NIST AI Resource Center. 
  • Cloud Security Alliance, “AAGATE: A NIST AI RMF-Aligned Governance Platform for Agentic AI,” December 2025. 
  • Cloud Security Alliance, “NIST AI Risk Management Framework: Agentic Profile.” 

Author

Related Articles

Back to top button