CloudAgenticAnnouncementsEnterprise AI

General Compute and Cerebras Team Up to Supercharge AI Coding with Wafer Sized AI Chips

The generative AI boom was built on chatbots, where users type a prompt and wait a few seconds for an answer. But as the industry races toward autonomous software agents, AI systems that plan, code, and execute tasks on their own, the underlying infrastructure is hitting a wall. The problem is no longer just intelligence; it is raw speed.

To break this bottleneck, infrastructure provider General Compute has signed a multi-year agreement to deploy Cerebras Systems‘ massive wafer-scale inference engines at scale. Going live in Q1 2027, the deployment is purpose-built for the most latency-sensitive workload in the enterprise AI space: agentic coding.

The Compounding Cost of Latency

An AI coding agent operates entirely differently than a standard conversational model. It does not just answer a question. It reads a repository, drafts a plan, writes code, runs tests, analyzes the failure logs, and revises its work in a continuous loop until the task is complete.

This process requires hundreds or even thousands of sequential inference calls. Because each step relies on the output of the previous one, latency does not average out across the workflow; it compounds into raw wall-clock time that a human developer spends waiting on the critical path.

The math is unforgiving. If an agent needs 400 model calls to finish a software task, standard hardware decoding at 100 tokens per second will take roughly 20 minutes to complete the job. At that speed, most developers abandon the workflow and switch contexts. But push that decode speed to 2,000 tokens per second, and the same 20-minute task shrinks to just 60 seconds, keeping the developer in the loop and fundamentally changing how the tool is used.

Bypassing the Memory Bottleneck

Hitting those extreme speeds requires a structural departure from traditional AI hardware. For a single user streaming tokens, the limiting factor isn’t usually compute power, it’s the time it takes to move data from external memory to the processor.

Cerebras solves this by building the largest chip in the world on purpose. The Cerebras Wafer-Scale Engine spreads model weights across massive on-chip SRAM the size of an entire silicon wafer. By keeping the model entirely on the chip, decode processes never have to wait on an external memory bus. This architectural advantage allows Cerebras to sit at the top of independent per-user speed leaderboards, pushing sub-second latency that standard GPUs struggle to match.

“In AI, speed is productivity,” said Sean Lie, CTO and co-founder of Cerebras. “An agent that takes hundreds of steps to finish a task is only as fast as its slowest step. Working with General Compute puts Cerebras speed in front of the developers building these agents, on a platform they already trust.”

Putting Wafer-Scale Power in Developers’ Hands

Despite the clear performance advantage, most AI software startups and enterprise IT departments cannot afford to buy and physically install a wafer-scale hardware system, nor do they want the operational headache of managing exotic silicon.

General Compute exists to bridge this gap. As a “neocloud,” the company purchases and manages the specialized hardware, adding Cerebras to its fleet as a dedicated tier for developers whose primary constraint is per-token latency. Customers consume the resulting computing power strictly as inference, backed by a single contract and set SLAs.

“The chips that win inference are not going to come from one vendor, and most customers cannot put a wafer-scale system on their own balance sheet,” said Finn Puklowski, co-founder and CEO of General Compute. “Agentic coding is where that speed is worth the most right now, so that is where we are starting.”

As the AI industry transitions from tools that simply draft text to agents that autonomously execute complex workflows, extreme inference speed is no longer just a luxury, it is a functional requirement. By making wafer-scale silicon accessible as a cloud service, General Compute and Cerebras are laying the tracks for the next wave of AI workers.

Read the full announcement here: https://www.generalcompute.com/blog/general-compute-cerebras-multi-year-agreement

Author:

Related Articles

Back to top button