Scaling AI inference

| Article

AI build-out has passed a threshold. For years, a large share of compute resources were devoted to training, driven by accelerator scarcity, frontier model races, and the need to assemble increasingly large clusters—often thousands to tens of thousands of graphics processing units (GPUs)—to train those models. By 2030, the AI workload mix is expected to shift toward inference, which will account for about 60 percent of AI demand, compared with about 40 percent for training (Exhibit 1). The industry is shifting from building AI intelligence to serving it at scale, a fundamentally different engineering challenge with a distinct cost structure (see sidebar, “Latency, throughput, and cost constraints vary across inference requests”).

Image description: A stacked bar chart shows demand by workload type under the continued momentum scenario. The total demand in 2026 was 120 gigawatts, of which AI training accounted for 29% of the total, AI inference 29%, and non-AI 42%. In 2030, the total demand is expected to be 278 gigawatts, with AI training making up 28% of the total, AI inference 42%, and non-AI 30%. AI inference is expected to make up 60% of all AI demand in 2030, with a per annum growth of 36%. Total AI growth is expected to be 30% per annum. End image description.

Indeed, C-suite leaders are no longer asking whether AI creates value but whether it does so at an acceptable cost to serve. Enterprises pay for AI use and infrastructure immediately. Process redesign, data integration, employee adoption, and measurable productivity gains take longer. An AI deployment can look very expensive before it looks successful, and executives expect costs to rise materially as autonomous systems take on an increasing share of work; much of AI adoption and the infrastructure demand required to support it may still lie ahead. But unless cost to serve decreases significantly, enterprises may not be able to adopt AI inference at scale.

To understand this new environment, McKinsey collaborated with the Global Semiconductor Alliance (GSA) to conduct interviews with AI architects from leading semiconductor players and hyperscalers. From these conversations, we’ve distilled five shifts that will define the era of AI inference:

  1. Enterprises will predominantly judge AI performance by cost per useful outcome rather than component benchmarks, potentially moving value toward the value chain layers that improve system economics.
  2. Workloads are becoming more diverse, creating demand for specialized silicon (which can happen through levers such as [semi-]custom silicon or chiplet designs).
  3. Compute for inference will be distributed across cloud, enterprise, edge, and device, with location dictated by economics, latency, and data control.
  4. Hardware—logic, memory, and networking—dictates what AI compute can deliver, while software dictates how much of that capability can be captured.
  5. Specialization pays only at scale, making open standards a necessity.

Below, we explore these shifts in detail, as well as propose how the semiconductor industry should respond to address the concerns of enterprise leaders and successfully scale AI inference.

1. Enterprises may weigh cost per outcome rather than generic benchmarks

The hardware road maps already point toward more performance, more memory, faster interconnect, and denser systems. The challenge is that each increment now costs more than the last, and the returns are nonlinear. So as enterprises decide where to invest in AI, they need better demand signals than generic benchmarks alone can provide.

Because enterprise demand is fragmented by industry, AI requirements should be derived from the business outcome. A retail recommender may need to serve millions of relatively low-cost requests with high availability and predictable unit economics, while an autonomous vehicle requires near-real-time inference with stringent latency and reliability requirements. These workloads have very different requirements for infrastructure, cost, performance, and service levels, which a single general-purpose benchmark cannot capture.

Instead of looking at whether the spec sheet improved, enterprises must consider what a given workload costs to serve reliably, including the cost per resolved service ticket, per line of accepted code, and per inspected unit. Given this shift, the companies building hardware and the customers deploying it must collaborate early in the process while architectural choices can still be shaped.

This change in what the industry measures will likely also move value. Today, value capture is heavily influenced by scarcity; leading accelerators, high-bandwidth memory, advanced packaging, and access to power all command premiums, partly because supply has struggled to keep pace. Durable value may accrue instead to the layers that measurably improve system economics, such as lowering cost, raising utilization, cutting latency, or increasing useful output per watt. A memory supplier that relieves a bandwidth bottleneck, an interconnect vendor that keeps a rack usefully coherent, a cooling or power-delivery player that lets density rise—each earns its position by improving what the whole system delivers.

Near-term gains will likely remain concentrated among memory suppliers, leading accelerator vendors, and a small group of custom-silicon and networking companies. But as constraints move beyond compute, power delivery, thermal management, connectivity, and sensing may offer opportunities for the many analog, microcontroller, and other fabless players still defining where they can differentiate. Specialization remains valuable when it solves a system-level problem rather than delivering a better line on a datasheet.

2. Diverse workloads call for specialized silicon

Multiple inference architectures will likely coexist as different workloads place materially different demands on AI infrastructure (Exhibit 2).

Image description: A table shows workload requirements and estimated share of 2030 AI inference token demand. Under AI training on the table, agentic, code, chat and search, writing, and machine to machine fall under the large language model (LLM) text and code category (82%); media generation and physical AI fall under the non-LLM category (13%); and voice makes up the remainder. Model types across all AI training categories include LLMs, multimodal models, diffusion models, world models, vision models, speech models, recommendation models, and reinforcement models.  Model scale is 10 to 100 billion parameters, up to 128,000-token context window, for agentic, code, and chat and search and is under 10 billion for writing, machine to machine, media generation, physical AI, and voice. Compute is low-latency throughput for agentic and chat and search; high throughput for code, writing, machine to machine, and media generation, and high performance or wattage for physical AI and voice. Memory or networking is high key-value-cache capacity for agentic; high capacity for code; low latency for chat and search, physical AI, and voice; high bandwidth for writing and machine to machine; and moderate memory capacity for media generation.  Serving pattern is as follows: for agentic, iterative agent loop and token by token with cache reuse; for code, long-form token by token with cache reuse; for chat and search, retrieve, then interactive token by token; for writing, token by token with cache reuse; for machine to machine, high-volume request and response; for media generation, long-form token by token with cache reuse; for physical AI, queued jobs, one at a time, 3–50 steps; and for voice, real-time bidirectional streaming. The workload requirement is heavier across model scale, compute, and memory for agentic code, and chat and search. It is moderate for model scale and heavier for compute and memory for writing and machine to machine. It is moderate across model scale and memory for media generation. Last, it is moderate across model scale and compute and heavier in memory for physical AI and voice. Note: Figures may not sum to 100%, because of rounding. Source: AMD; DeepSeek; Google; JEDEC; Lawrence Berkeley National Laboratory; NVIDIA; OpenAI; UALink Consortium; Uptime Institute; McKinsey analysis End image description.

Much of today’s installed base was designed for training and is now being used for inference. That creates a mismatch, but not the one spec sheets suggest. Training rewards raw compute; decode rewards memory bandwidth. Generating each token means streaming the model’s weights and accumulated KV cache out of memory to do comparatively little arithmetic. So bandwidth rather than peak FLOPS1 sets how fast a node can serve, and peak compute has been growing roughly twice as fast as memory bandwidth for two decades.

The GPU-plus-HBM2 configuration that dominates deployments today will continue to serve a large share of workloads, but it is unlikely to remain the single default as inference requirements become more specific. The more likely outcome is a wider spread of configurations: nodes weighted toward bandwidth or toward capacity depending on the workload, HBM generations that widen the interface alongside custom base dies tuned to a particular buyer’s needs, advanced packaging that shortens the distance data travels, cheaper memory tiers holding the colder parts of the cache, and SRAM-heavy designs attacking the bottleneck from the opposite end. Which balance prevails is unsettled and will likely differ by workload rather than converge.

Consider a compliance tool that reads a 300-page contract and returns a short list of flagged clauses and a voice assistant that streams a spoken reply token by token. The first is prefill-heavy and compute-bound; the second is decode-heavy and bandwidth-bound. Run on the same hardware, each pays for capabilities it does not use.

The same split appears inside a single request. Prefill, where the model processes the prompt, is compute-intensive. Decode, where it generates tokens one at a time, depends more on memory bandwidth. Long inputs increase prefill work; long outputs increase decode work. So a text agent and a voice agent performing the same task may require different hardware balances, and different latency targets can favor aggregated or disaggregated serving.

The industry is now designing around this. Hyperscalers and AI labs are building custom accelerators for their highest-volume workloads. SRAM-based players including Groq, Cerebras, SambaNova, and d-Matrix target the bandwidth-bound half of the problem. And a wave of pairings has emerged to split phases across purpose-built silicon: NVIDIA pursued a licensing agreement and talent transfer with Groq, allowing it to absorb Groq’s purpose-built inference technology; AWS paired Trainium prefill with Cerebras decode; AMD partnered with Cerebras; and Intel and SambaNova introduced an inference blueprint routing prefill to GPUs and decode to reconfigurable dataflow units (RDUs).3 OpenAI’s Jalapeño takes the opposite bet: a single balanced accelerator that activates the right mix of compute, memory, and network per phase, keeping model state local rather than moving it between hardware pools.4 In addition, many players are exploring chiplet-based architectures to provide customized solutions on a standard compute platform.

There’s one main risk with specialization: Chip cycles run in years while model architectures change in months, so purpose-built silicon can arrive optimized for workloads that have already moved. Compressing those cycles is now an explicit industry objective. Reinforcement-learning floorplanning has been in production use across several generations of Google’s tensor processing units, and Jeff Dean has argued publicly that AI methods could bring design timelines from 18 to 30 months down to three to six.5 OpenAI’s Jalapeño is the clearest proof point so far: co-developed with Broadcom, it moved from first RTL to manufacturing tape-out in nine months—compared with 18 to 24 months for a typical ASIC designed from scratch. OpenAI used its own models to explore implementations, compress verification loops, and optimize the chip’s arithmetic circuits. Machine-learning-based placement and optimization are now standard features in the major electronic-design-automation flows rather than experimental add-ons, and a set of well-funded start-ups is targeting the rest of the design flow. If these efforts deliver even part of what they promise, teams could design for the workloads they actually have rather than the ones they predict, and workload-specific silicon would become far more viable.

3. Where inference runs is based on economics, latency, and data control, not technology

Compute for inference is likely to be increasingly distributed across cloud, enterprise, edge, and device. The right location will vary by workload, shaped by data location, economics, latency, reliability, and control. Technical feasibility is rarely a barrier; it’s largely possible to run inference in each location. What differs is who is willing to pay for it (Exhibit 3).

Image description: A table shows the technical readiness and commercial clarity of inference compute location. For the cloud location (for example, frontier model serving) and for the enterprise or co-located location (for example, bank scoring fraud on-premises), they are technically ready and their commercial clarity is settled. For the edge location (for example, visual inspection on the line), it’s technically ready but clarity is unresolved. Last, for the device location (for example, on-phone message summarization), its technical readiness is emerging or use-case dependent and its clarity is partial. End image description.

Cloud. Cloud will likely remain central for frontier models, elastic workloads, rapid model refreshes, and applications where data is already centralized. But it may no longer be the default for every workload as economics and other requirements become more important.

Enterprise. Enterprise and co-location infrastructure may grow where governance and sovereignty are critical. For example, banks scoring fraud against customer records will likely want to keep sensitive customer records on their premises. These workloads still need data-center-class systems; they simply sit in private clusters rather than hyperscale ones.

Edge. Edge serves workloads needing substantial compute near where data is generated—such as a plant running visual inspection on every unit coming off the line, where moving the video costs more than processing it locally. While edge is technically feasible, no one has settled who pays. Telecom offers a cautionary tale: Operators funded billions in 4G, consumers paid a modest monthly subscription, and much of the value accrued to the applications built on top. Telcos, utilities, and site operators will hesitate to move to edge until the payment question is clear.

Device. This deployment option wins when the model fits the endpoint’s memory, power, and thermal envelope and when immediacy, privacy, or offline operation matters. For example, a phone summarizes messages locally and escalates only the harder requests upstream.

4. Hardware dictates what AI compute can deliver, while software dictates how much can be captured

Accelerators will remain central going forward, but they may no longer determine inference economics on their own. Within silicon, performance is increasingly shaped by three interdependent elements: Logic performs the compute, memory feeds it, and networking connects and scales it. Looking ahead, memory bandwidth and networking are likely to become increasingly important constraints—leading some in the industry to argue that the network is becoming part of the computer itself.

That ceiling is physical. As context windows lengthen and agentic workloads multiply KV-cache traffic, memory capacity and bandwidth set how many concurrent requests a node can serve. Networking sets how far a single request can usefully be spread across a rack before latency erodes the gain. Power delivery and thermal management (for example, liquid cooling and emerging approaches such as microfluidics) increasingly gate density before compute does. Many of the latest semiconductor innovations (such as silicon photonics or advanced packaging) attack the same problem: the energy and latency cost of moving data rather than processing it.

The near-term cost curve, though, might be won in software. A coding assistant that rereads the same repository on every request can have much of its prefill cost removed through cache reuse. A customer service fleet that routes simple queries to a small model and escalates only the more complex ones changes cost per resolution without touching the hardware. Batching, scheduling, quantization, and data placement behave the same way: Gains compound, and they make existing hardware less expensive to run.

The implication is that inference infrastructure is a full-stack optimization problem (Exhibit 4). Hardware defines what a system can deliver, while software (and ultimately enterprise data and workflow integration) determines what it actually delivers.

A triangle diagram, in which compute, software, and memory and networking make up the three sides of the triangle, shows inference optimization techniques across the three sides. Under compute, techniques are as follows: Custom silicon: Design specialized chips tuned to inference workloads. Distributed computing: Orchestrate model execution (prefill versus decode) across differentiated infrastructure. Advanced nodes (less than 3 nanometers): Achieve energy efficiency and density. Advanced packaging (2.5D or 3D and Chip-on-Wafer-on-Substrate.): Achieve higher bandwidth and chiplet scaling Under software, techniques are as follows: Model optimization techniques (for example, quantization, mixture of experts, distillation, KV-cache optimization) to cut compute and memory cost. Software scheduling techniques (for example, continuous batching, topology-aware scheduling) Last, under memory and network, techniques are as follows: Scale-up interconnect: Use high-speed interconnects to minimize intra-node communication latency, Optical interconnect evolution: Deploy optics-integrated switches (copackaged optics) for higher bandwidth density and lower power per bit, Compute Express Link and dynamic random-access memory tiering: Pool and offload key-value (KV) data to cheaper, shared memory tiers, Low-power double data rate: Integrate low-power memory for energy-efficient inference chips End image description.

5. Specialization only pays at scale, making open standards a necessity

As AI becomes embedded in critical workflows, governments and enterprises are placing greater weight on where data is processed, which models are used, where infrastructure resides, and who can access the outputs. For a growing set of workloads, sovereignty, resilience, and regulatory requirements have become as important as cost and performance in deployment decisions.

The effect is that requirements are diverging rather than converging. A national health service, a defense ministry, and a regional bank may all want the same capability but under different constraints on data residency, approved suppliers, model provenance, and local ecosystem participation. Demand that once converged on a handful of hyperscale regions could instead resolve into dozens of in-country deployments, each with its own specification. For semiconductor companies, that is both a demand driver and a source of complexity, involving incremental build-out on fragmenting requirements.

Sovereignty is not the only force that could pull in this direction. Ordinary competition will have a similar effect: Hyperscalers optimizing around their own accelerators and software, merchant vendors defending full-stack platforms, and robotics and automotive players designing to their own latency, safety, and power envelopes all have reasons to differentiate rather than converge. In both cases, the likely result is several viable inference stacks. Sovereignty layers additional constraints on top of the divergence that competition already creates.

Fragmentation is where open models and standards become economically important rather than ideological. Open-weight models let an operator take the same capability into a sovereign cloud, a private cluster, or an edge deployment and tune it to whatever silicon is locally available. Closed models can lead on capability and managed integration, but they concentrate more of the stack within one provider’s ecosystem, which can be harder to accommodate when customers require control over deployment, data, or underlying infrastructure. The eventual balance between open and closed models will shape how diverse the inference market remains.

Open interconnect and software standards do the same work on the hardware side. A specialized architecture only earns its cost at volume; if serving it in fifteen jurisdictions means fifteen software stacks, the integration burden pushes buyers back toward general-purpose platforms that are easier to replicate anywhere. Common standards allow providers to adapt specialized architectures to local requirements without building a separate software stack for each jurisdiction, preserving the scale that makes specialization viable.


Inference is turning AI from a race to build intelligence into a race to serve it economically at scale. The five shifts described here point the same direction: Infrastructure is reorganizing around the value of each workload rather than the capability of each component.

That is why no single company—and no single layer—can solve inference economics alone. Cost to serve is set by how silicon, memory, networking, software, power, cooling, and data architecture work together, and most of those decisions are expensive to reverse once made. Industry alliances and partnerships can help the ecosystem get them right the first time, through common standards, interoperable architectures, shared road maps, and commercial models that align the companies designing infrastructure with those funding and deploying it.

Serving intelligence economically can turn AI from a capability a few can afford into infrastructure that entire economies can build on.

Explore a career with us