AI Business Strategy

Why AI Is Only as Good as the Network Behind It

By Chris Carreiro, CTO, Park Place Technologies

To answer this question, most will consider surface-level factors like GPUs purchased, racks deployed, and raw compute power available. While these are important considerations, the ready-or-not answer lies deeper within your technical estate.  

Compute has skyrocketed in the past few years, but the infrastructure that supports it often gets left on the launch pad. For enterprise AI to succeed, organisations must look beyond GPUs and determine if their network is ready to handle AI inferencing.  

Overcoming GPU blind spots 

AI performance hinges on infrastructure, not raw compute power. In the rush to accelerate AI, many organisations develop a GPU “blind spot” and overlook crucial infrastructure issues that may be holding them back.  

Working around these GPU blind spots starts by understanding AI training and inferencing. These workloads depend on a chain connecting storage, networking, servers, GPUs, and front-end applications. The slowest link in this chain sets the pace for everything else.  

An organisation could have the most powerful GPUs on the market, but without the storage or networking infrastructure to match, its investment will fall flat. This is why network observability is becoming a critical component of AI operations. Understanding the health of data paths delivers more value than measuring the performance of individual components. 

Utilisation, not capacity, determines return. Before purchasing GPUs, organisations should consider their network observability and complete any necessary upgrades. This is where the most avoidable losses occur.   

How GPU clusters break tradition  

Traditional data centre traffic is “north/south,” originating from a user or system. In general, a request comes in, hits a server, and a response is sent back up the chain. This pattern is volume-based, meaning the risk of a slowdown increases as user requests increase. In GPU environments, data is exchanged in “east/west” patterns from server to server and rarely involves direct user interaction. 

Enterprise networks are designed for north/south traffic, which can tolerate some latency without degrading the user experience. AI-dense east/west traffic, on the other hand, is latency-sensitive. GPU clusters can quickly become bogged down by network congestion and inconsistent packet delivery. This latency slows down the entire cluster, even if it occurs on only one GPU, turning a massive investment into a major disappointment. 

A failure of this magnitude rarely occurs as a hard outage. It’s a degradation that can occur over months before an organisation even recognises there’s an issue. The business case for GPU investments quietly erodes as a result.  

Why device-level monitoring falls short 

Enterprise data centre health has long been determined through individual device reporting. For example, a network management platform checks each device in the server room every few minutes to determine CPU and memory usage, chassis temperature, error frequency, and more. This device-level monitoring helps organisations quickly identify hard failures and downward trends that may warrant an upgrade.  

GPU server environments are a different story because health is a property of the data path, not the device. Every hardware component can appear healthy while the workload underperforms. This is because the estate is spread across network, storage, compute, and hardware silos, each with its own monitoring. So, instead of using software to track device health across the server room, organisations must investigate each silo to diagnose issues. 

The congestion that slows AI down can occur in millionths of a second. Even if an organisation could use device-level monitoring to track GPU health, it wouldn’t be fast enough to detect a problem. It’s like checking motorway traffic once per hour. At that pace of monitoring, it may seem like traffic flows freely. But traffic jams could have formed and cleared between checks.   

Protecting GPU investments with observability 

The key to strengthening the AI business case is network observability. 

Organisations need to dig deeper and analyse data paths rather than devices. Consider how data moves back and forth through the environment and watch for slowdowns. Monitor signs of degradation, not just failure, to address latency issues as early as possible. This can help uphold the GPU business case by reducing the time it takes to identify and repair any issues.  

Network observability is often the difference between reacting to problems and preventing them. Organisations that can see traffic patterns, bottlenecks, and latency trends across the entire data path are better positioned to maintain performance as AI workloads scale. 

Another benefit is evidence-based capacity planning.  

Instead of purchasing more GPUs to address speed concerns, organisations can use network observability to determine the root cause of a slowdown. This, too, could save the organisation time and money. Remember, a bandwidth bottleneck isn’t always a GPU capacity issue. The true problem could be hidden within the network itself. 

Ready or not? 

Determining AI readiness isn’t a matter of asking how many GPUs an organisation has in the data centre. It’s whether the network surrounding those GPUs can keep up, and what measures are in place to respond if the answer is “no.”  

Related Articles

Back to top button