Business6 min read

The Hidden Bottleneck Behind Nvidia’s New AI Chips

Why HBM memory and advanced packaging matter to AI chip supply. A plain-English guide to TSMC CoWoS, Nvidia Rubin, and the manufacturing steps after lithography.

Loading the player connects to YouTube (a Google service) and plays the video in YouTube's privacy-enhanced mode. Watch on YouTube instead without loading the player here.

An AI chip can have an impressive design and still be difficult to deliver at scale. Manufacturing the compute dies is only part of the job: those dies must work with memory and interconnects in a complete package.

Advanced packaging is the manufacturing step that brings these pieces together. It can become a supply constraint even when a company has secured wafer capacity. That makes the AI chip race partly a contest over the ability to assemble a working product, rather than simply design faster processors.

Our previous article, The $400 Million Machine That Could Decide the AI Chip Race, explained lithography. This companion to Video #8 follows what happens after the circuits are printed.

Why a GPU needs more than compute

Nvidia’s July 2026 architecture description lists 336 billion transistors, up to 288 GB of HBM4 memory, and up to 22 TB/s of memory bandwidth for Rubin. It also describes two compute dies connected within one package through Nvidia’s High-Bandwidth Interface (see Sources, [1]).

Those figures describe different capabilities. Transistors support computation. Memory capacity determines how much data can be held nearby. Bandwidth describes how quickly data can move.

More arithmetic capability is of limited help when the next operation is waiting for its data. Nvidia identifies memory movement as a particular constraint during the token-generation phase of inference (see Sources, [1]).

This gives readers a useful way to interpret accelerator specifications: ask about compute, memory capacity, and data movement together. A larger number in one column does not automatically remove a limitation in another.

What HBM brings to the package

HBM means high-bandwidth memory. It uses stacked memory dies, integrated close to the processor rather than treated as a distant store of data. TSMC describes CoWoS as integrating compute chips with HBM stacks (see Sources, [2]).

Think of the compute units as a workshop and memory as its supply room. Adding more workers helps only if materials can reach them quickly enough. The layout and connections between the two become part of the production system.

This analogy is deliberately limited: a chip package has electrical, thermal, and physical requirements that a room does not. Its value is to show why memory and its connections deserve attention alongside the processor.

What CoWoS actually does

CoWoS stands for Chip-on-Wafer-on-Substrate. TSMC describes it as a family of advanced 2.5D packaging technologies that integrates compute chips and memory using dense interconnections (see Sources, [2]).

An interposer provides a connection structure between the components. CoWoS-S uses a silicon interposer; CoWoS-R uses a redistribution-layer interposer. CoWoS-L combines a redistribution structure with local silicon interconnects. These are different approaches within the same technology family (see Sources, [2]).

“Packaging” can sound like a protective box added at the end. Here, it helps define how the product’s components communicate. Choosing the integration approach is therefore part of engineering the accelerator.

TSMC lists CoWoS-S interposers up to roughly 3.3 times reticle size and recommends CoWoS-L or CoWoS-R for larger sizes. It says its first 3.5-times-reticle CoWoS-L entered volume production in 2024 (see Sources, [2]).

These figures describe interposer area relative to a lithography exposure field. They do not mean that an individual compute die is printed in one exposure at several times that field size.

Why wafer capacity is not enough

Consider a simplified manufacturing chain: compute dies, memory, packaging, boards, servers, and deployment. Each stage must supply compatible components at the required time.

As an illustrative example, imagine enough compute dies for 100 accelerators but enough compatible memory and packaging for only 70. The remaining dies cannot become 30 additional complete accelerators merely because the wafer manufacturing step is finished. These numbers are an explanation, not reported production data.

TrendForce’s April 30, 2026 industry report described constraints in both leading-edge wafer production and advanced packaging. It said the pressure also extended to equipment, substrates, and packaging materials (see Sources, [3]).

The report said Nvidia had secured capacity and materials across several stages, including wafers, CoWoS, HBM, substrates, and PCBs. That is a reported supply-chain strategy, rather than proof that every potential shortage has been eliminated (see Sources, [3]).

A bottleneck can move

Packaging should not be presented as the permanent, exclusive limit on AI growth. Capacity investment, product redesigns, demand changes, and alternative integration approaches can change where the tightest constraint sits.

TrendForce’s April report expected some easing of the global 2.5D packaging shortage in 2027 as capacity expanded. That was a forecast, not an observed outcome or a guarantee (see Sources, [3]).

For a business following this market, the useful questions are therefore specific: which technology, which product generation, which supplier, and which delivery period? “Packaging capacity” is too broad a phrase to answer all four.

Adding nominal capacity also does not automatically establish how many qualified, working products can be shipped. A forecast should be read with its assumptions and date attached.

The supply chain behind a finished system

Nvidia said in May 2026 that Vera Rubin’s ramp involved more than 350 factories across 30 countries. The release described platform production and scheduled production shipments to begin in the fall; the announcement itself was not evidence that every customer had received systems (see Sources, [4]).

That scale illustrates the coordination involved. An accelerator eventually becomes part of a server and a rack, and those systems need a site that can power, cool, and operate them.

The Hidden Infrastructure Companies Behind the AI Boom follows that wider infrastructure chain. This article focuses further upstream, on assembling the semiconductor product that those facilities will use.

What this changes about the AI chip race

The business implication is an inference from the manufacturing chain: a stronger design is commercially valuable only when it can become a complete, reliable product on a useful schedule.

A supplier should therefore be judged on more than an announced specification. Its delivery capability depends on compatible manufacturing stages working together. Equally, a shortage headline does not establish that every packaging company will benefit or that any particular investment will earn a return.

Lithography remains essential. Memory and advanced packaging add further dependencies. Power and cooling matter after the accelerator is assembled.

The question behind Video #8 is practical: can the industry manufacture and deliver the complete system quickly enough to turn the promise of a new AI chip into usable computing capacity?

Sources & References

  1. [1]Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI — NVIDIA
  2. [2]CoWoS advanced packaging technology — TSMC
  3. [3]AI Competition Turns into a Supply Chain Arms Race, Tightening Advanced Packaging and 3nm Capacity — TrendForce
  4. [4]NVIDIA Vera Rubin Ramps Into Full Production to Power Agentic AI Factories Worldwide — NVIDIA