Article
← All posts
ai

The Last Millimeter

Reading time: 7 minutes

TL;DR

  • An AI accelerator is not one chip. It is a sandwich: a logic die and stacks of high bandwidth memory bonded onto a silicon interposer. That final assembly step is called advanced packaging, and TSMC’s version, CoWoS, dominates it
  • Fabs can now make the silicon faster than anyone can package it. Per Digitimes and industry analysts, packaging, not wafer fabrication, is the binding constraint on AI accelerator production in 2026
  • TSMC’s CEO told shareholders in June 2026 that CoWoS capacity is extremely tight and sold out through 2026. Analyst tracking puts lead times at 52 to 78 weeks, with backend facilities booked into 2027
  • TSMC is scaling CoWoS from roughly 13,000 wafers per month at end 2023 to a targeted 125,000 to 130,000 by end 2026, a near 10x expansion. It’s still not enough
  • Per CNBC, Nvidia has reserved the majority of CoWoS capacity through at least 2027, leaving everyone else to fight for the remainder
  • Packaging prices are rising 2 to 4 times faster than wafer prices. The scarcity is not in the silicon. It’s in the glue

1. Your GPU Is a Sandwich

The public story of chipmaking is transistor shrink: 5 nanometers, 3 nanometers, now 2. That story is doing fine. Foundries fabricate leading-edge logic dies at healthy yields.

The problem is what happens after. A modern AI accelerator only works because a logic die sits millimeters away from stacks of high bandwidth memory, wired together by tens of thousands of microscopic connections that ordinary circuit boards cannot support. Building that assembly is a manufacturing discipline of its own called advanced packaging, and for high-end AI chips the dominant process is TSMC’s CoWoS, short for Chip on Wafer on Substrate.

Inside the Package: Where the Bottleneck Lives

  CROSS-SECTION OF AN AI ACCELERATOR
  ─────────────────────────────────────────────

   [HBM] [HBM]  [ LOGIC DIE ]  [HBM] [HBM]
     │     │    (the "GPU")      │     │
  ═══╪═════╪══════════╪══════════╪═════╪═══
  [        SILICON INTERPOSER            ] ◄── THE
     tens of thousands of micro-wires        CHOKEPOINT
  ═══════════════╪═══════════════════════
  [           PACKAGE SUBSTRATE          ]
  ═══════════════╪═══════════════════════
  [          CIRCUIT BOARD               ]

  The fab makes the logic die.
  The memory makers supply the HBM.
  CoWoS bonds it all onto the interposer.
  Miss any layer and you own an expensive
  pile of perfectly good, useless silicon.

Without that bonding step, a flawless 3 nanometer wafer cannot become a functional AI chip. It’s the last mandatory station on the assembly line, and TSMC controls nearly every workbench that can run it at commercial scale.

2. Sold Out, Again

The supply numbers read like the memory story from Article 1, one step further down the chain.

At the end of 2023, TSMC’s CoWoS capacity was roughly 13,000 wafer starts per month. By the end of 2024 it was about 35,000, by the end of 2025 roughly 75,000, and the target for the end of 2026 is 125,000 to 130,000, per Silicon Analysts tracking. That’s a near tenfold expansion in three years, one of the most aggressive capacity ramp ups in semiconductor history.

And it’s still not enough. TSMC CEO C.C. Wei told the company’s annual shareholder meeting on June 4, 2026 that CoWoS capacity remains extremely tight and sold out through the year. Nvidia’s management has said in earnings calls that assembly capacity is oversubscribed through at least mid 2026. Silicon Analysts reports both CoWoS variants fully booked with lead times of 52 to 78 weeks and backend facilities effectively sold out into 2027. And per an April 2026 CNBC report, Nvidia has reserved the majority of TSMC’s CoWoS capacity through at least 2027, with that capacity growing at roughly 80% per year and still falling short of demand.

The clearest signal is price. Advanced packaging prices are rising 2 to 4 times faster than the wafers themselves. When the assembly step inflates faster than the silicon inside it, the market is telling you exactly where the constraint moved.

3. Why Gluing Is Hard

Calling it glue undersells the physics, which is precisely why capacity cannot simply be willed into existence.

The interposer is itself a piece of silicon, often larger than the chips it carries, patterned with interconnects far finer than any circuit board. As AI chips grow, packages now span multiple logic dies and eight or more HBM stacks, and the newest CoWoS variant used for Nvidia’s flagship parts stitches local silicon bridges into a redistribution layer. Larger packages mean new failure modes, including warpage, the literal bending of the assembly under thermal stress, which can crack thousands of connections at once.

The tooling compounds the problem. Thermal compression bonders and ultra-precise pick-and-place machines carry 12 to 18 month lead times of their own, per supply chain reporting. The bottleneck has a bottleneck: you cannot expand packaging lines faster than the equipment makers can build the equipment.

This is also why the industry’s famous logistics quirk exists: analysts call it the golden screw problem. A $400,000 AI server cannot ship while its accelerator sits unpackaged, so OEMs hoard every peripheral component and wait. One missing assembly step idles inventory across the entire bill of materials.

4. Who Gets Squeezed

When one company controls the chokepoint and one customer reserves most of it, the math for everyone else gets ugly.

Broadcom, AMD, and the custom silicon programs at hyperscalers all draw from the same CoWoS pool. Broadcom has publicly flagged TSMC advanced node and packaging capacity as the limiter on its 2026 AI chip supply. TSMC has begun routing overflow to outsourced assembly partners ASE and Amkor, whose advanced packaging revenues are growing accordingly, but qualifying a second source for a bleeding-edge package is itself a multi-quarter engineering project.

There is also a policy wrinkle worth watching. Current US tariffs target manufactured goods, and CoWoS is performed as a manufacturing service in Taiwan, so it has so far escaped direct tariff exposure. Analysts flag a credible scenario in which advanced packaging is reclassified as a derivative product in future reviews, which would raise the landed cost of US-delivered accelerators overnight.

5. The Escape Routes

Three paths out are visible, all measured in years.

More of the same, faster. TSMC’s newest dedicated packaging facilities, including the AP7 site planned for up to eight production buildings, plus the ASE and Amkor overflow, are the near-term relief valve. This is the 2026 to 2027 story.

Bigger canvases. Panel-level packaging replaces round 12-inch wafers with rectangular panels, lifting area utilization from about 57% to over 80%, essential as packages outgrow the wafer itself. Glass substrates, flatter and more thermally stable than today’s organic materials, are expected in first commercial applications around late 2027.

Geographic spread. Advanced packaging is finally being built outside Taiwan: Amkor in Arizona, and SK Hynix’s $15 billion Indiana facility for HBM packaging, which per company announcements does not reach mass production until 2028. The pattern by now should be familiar from memory fabs and transformers: the relief is real, and it’s two years away.

The Bottom Line

The AI supply chain spent two years learning that the constraint is never where the headlines are. First it was GPUs, then memory, then power. In 2026 the tightest link is a step most people have never heard of: bonding finished chips onto a silicon interposer, a process one company dominates, one customer has mostly reserved, and whose prices are inflating faster than the silicon it assembles.

The transistors will keep shrinking. The packages will keep growing. And until the panel lines and the Indiana and Arizona plants come online, every AI roadmap on Earth runs through the last millimeter, where the world’s most advanced technology waits its turn for very hard gluing. The encouraging part is hiding in the same numbers: a near tenfold capacity ramp in three years is one of the fastest industrial responses ever mounted to a bottleneck, and the escape routes are funded and under construction. The last millimeter is a queue, not a wall.


Sources

  1. TSMC Annual Shareholder Meeting remarks, C.C. Wei, June 4, 2026 (CoWoS extremely tight, sold out through 2026), via industry coverage
  2. Silicon Analysts, CoWoS capacity tracker and Q1 2026 foundry allocation analysis (13,000 wpm end 2023 to 125,000-130,000 wpm targeted end 2026; 52 to 78 week lead times; backend sold out through 2027; packaging prices rising 2 to 4x wafer prices)
  3. CNBC, April 8, 2026 (Nvidia reserving majority of CoWoS capacity through at least 2027; ~80% annual capacity growth still short of demand)
  4. Digitimes briefing, April 10, 2026 (advanced packaging, not wafer fabrication, as the binding constraint on AI accelerator production)
  5. Nvidia earnings call commentary (CoWoS assembly capacity oversubscribed through at least mid 2026)
  6. SupplyICs market intelligence, May 2026 (12 to 18 month packaging equipment lead times; golden screw effect on OEM inventories; ASE, Amkor, Intel Foundry expansion)
  7. Reuters and industry coverage of Broadcom remarks, March 2026 (TSMC advanced node and CoWoS capacity as 2026 supply limiter)
  8. TokenRing / FinancialContent analysis, January 2026 (AP7 facility plans; CoWoS-L warpage challenges; panel-level packaging 57% to 80%+ utilization; glass substrates late 2027)
  9. NextWaves Insight, April 2026 (tariff treatment of packaging as a manufacturing service; reclassification risk)
  10. SK Hynix company announcements (Indiana advanced packaging facility, ~$15B, mass production targeted 2028); Amkor Arizona facility announcements