The HBM4 shortage is being covered as a procurement story: who signed which multi-year agreement, which hyperscaler locked what allocation. That framing misses the mechanism. High-bandwidth memory is scarce because inference consumes it structurally โ every long-context agent is a memory-resident process whose footprint scales linearly with context, and every percentage point of HBM you cannot buy is agents you cannot run. This guide does the arithmetic end to end: bytes per token, agents per GB, wafers per bit, dollars per stack. The conclusion is uncomfortable for the "AI demand is infinite, supply will catch up" narrative: the binding constraint is physics of the KV cache meeting a wafer allocation that is structurally rationed.
The one formula, reused
From our KV-cache guide: KV bytes = 2 ร layers ร KV heads ร head dim ร bytes/value ร tokens. The four architectures this article uses, site-verified:
- MHA (e.g. no KV sharing, 64 layers, 64 KV heads, 128 dim, fp16): 2 ร 64 ร 64 ร 128 ร 2 B = 2,097,152 B = 2 MiB/token
- GQA-8 (Qwen3-32B: 64 layers, 8 KV heads): 262,144 B = 256 KiB/token
- MQA (1 KV head): 32,768 B = 32 KiB/token
- MLA (DeepSeek-V3: 576 latent elements ร 61 layers ร 2 B): 70,272 B โ 68.6 KiB/token
Agents per GB: the only capacity number that matters
A concurrent agent is its KV cache resident in HBM for the session's lifetime. Divide usable HBM by the per-session footprint and you get the only number procurement actually buys.
Session footprints and agents per GB (1024-based KiB/MiB, GB = 10^9 B):
| Architecture | B/token | 100k ctx / session | Agents per GB at 100k | 1M ctx / session | Agents per GB at 1M |
|---|---|---|---|---|---|
| MHA (fp16) | 2,097,152 | 209.7 GB | 0.0048 | 2,097 GB | 0.00047 |
| GQA-8 (fp16) | 262,144 | 26.2 GB | 0.038 | 262.1 GB | 0.0038 |
| MLA (bf16) | 70,272 | 7.03 GB | 0.142 | 70.3 GB | 0.0142 |
| MQA (fp16) | 32,768 | 3.28 GB | 0.305 | 32.8 GB | 0.0305 |
Derived directly: at 100k context, 1 GB of HBM serves 0.14 MLA agents or 0.038 GQA agents. One 288 GB GPU's raw capacity holds only ~1,100 concurrent 8k-context sessions even under this article's most compact architecture (MQA fp16, ~262 MB per session) โ ~137 with GQA-8; at agentic 100k context the same silicon holds 35 MLA agents or 9.5 GQA agents. The compute is idle long before the memory is full โ decode is memory-bandwidth-bound and the cache is memory-capacity-bound.
The wafer mechanism: every HBM stack is ~3 DDR5 modules not built
At Hot Chips 2026, Micron quantified what the industry muttered for years: for the same number of bits, HBM consumes roughly 3ร the wafer area of DDR5 โ and the penalty is widening with every generation1. The causes stack: TSV-thinned dies sacrifice usable wafer edge, each of a 12-Hi stack's dies carries routing overhead for vertical buses, base-logic dies eat leading-edge capacity, and stacking yield multiplies per-die yields. The consequence is arithmetic, not sentiment: a wafer start diverted to HBM produces one third the bits. Every 288 GB HBM-equipped GPU implicitly forgoes roughly 864 GB of DDR5-class supply from the same wafer input.
TrendForce puts HBM wafer input at ~18% of total DRAM wafer input by end-2025, rising to ~30% by end-2027 โ but HBM as a share of DRAM bit supply only reaches ~13% in the same window2. That gap โ 30% of wafers, 13% of bits โ is the ~3ร area penalty visible in industry statistics. This is why "build more fabs" is not a 2026 answer, and why Samsung's and SK hynix's combined DRAM inventories falling under ten days of supply in early September 20263 is a rationing signal, not a demand signal: the allocation mechanism, not the price mechanism, clears this market.
Rubin Ultra: the 288-to-192 GB step-down, recomputed
Rubin's base GPU ships 288 GB of HBM4. Rubin Ultra was announced at GTC toward 1 TB of HBM4E across 16 stacks in a 12-Hi configuration โ and by September 2026, Nvidia is reportedly testing variants at 192 GB and 256 GB, some stepping back from HBM4E to HBM4 and from 12-Hi to 8-Hi stacks45. Nvidia's public line is "roadmap intact"; the samples say the memory budget is not. TrendForce attributes the retreat to DRAM tightness through 2027 and uncertain 12-Hi HBM4E validation โ a supply-side call, not a design preference. The 8-Hi swap is itself the wafer math again: shorter stacks mean each fixed pool of known-good DRAM dies yields more complete stacks, trading capacity per stack for system output.
That 288 โ 192 GB cut is exactly 33% less HBM per GPU. Recompute the agent counts (budget = capacity โ 40 GB reserved for weights, activations, and runtime on a ~100 GB-class serving model with quantized weights; smaller model reservations only shrink the gap's cause, the slope is linear in capacity):
| GPU config | Free KV budget | MLA agents @ 100k | GQA agents @ 100k | MLA agents @ 1M | GQA agents @ 1M |
|---|---|---|---|---|---|
| 288 GB (Rubin) | 248 GB | 35.3 | 9.5 | 3.53 | 0.95 |
| 256 GB (Rubin Ultra tested) | 216 GB | 30.7 | 8.2 | 3.07 | 0.82 |
| 192 GB (Rubin Ultra tested) | 152 GB | 21.6 | 5.8 | 2.16 | 0.58 |
At 192 GB, a flagship 2027 accelerator serves 39% fewer 100k-context MLA agents than the 288 GB part it succeeds (21.6 vs 35.3) โ and GQA-fleet operators drop from 9.5 to 5.8 concurrent agents per GPU. Peak FLOPS staying constant is irrelevant: the constraint that binds is residency, and residency was cut by a third.
Price and BOM: memory is now a plurality of the machine
The spot-vs-contract spread is the cleanest supply signal available. Early September 2026: a 36 GB HBM3E stack trades at roughly $2,100 spot vs $300โ$400 under long-term agreements67 โ a 5โ7ร premium. Large buyers do not pay spot, and that is the point: suppliers have allocated the large majority of 2027 capacity to multi-year contracts (Seoul Economic Daily reports ~70% of output locked through as far as 2031 among Microsoft, Nvidia, and Google)8. The spot price is what the unallocated remainder costs. Contract DRAM tells the same story in slower motion: TrendForce's February 2026 revision put conventional DRAM contract prices at +90โ95% QoQ in 1Q26 โ note that is a near-doubling in one quarter, compounding with Q2's +58โ63% to roughly 3ร in six months9 โ and 2027 HBM4 contract negotiations are expected to lift prices "severalfold"2.
What that does to a GPU's bill of materials:
| Line | Estimate | Source type |
|---|---|---|
| H100: HBM โ $1,350 of โ $3,320 build cost (~41%) | per-stack LTA math | BOM teardown10 |
| B200: 8 ร 24 GB HBM3E โ $2,900 of โ $6,400 (~45%) | per-stack LTA math | BOM teardown10 |
| B300 (288 GB, 8 ร 36 GB): $2,400 at $300/stack โ or $16,800 at spot pricing | 8 ร $300 vs 8 ร $2,100 | computed6 |
| GB300 rack: memory โ $374k of โ $3.99M (~9.4%) | Morgan Stanley BOM | rack estimate11 |
| VR200 rack: memory โ $2.0M of โ $7.8M (~25.7%) | Morgan Stanley BOM | rack estimate11 |
A generation ago the logic die was the GPU's cost; on Blackwell-class parts HBM is approaching half the manufacturing cost, and at rack scale memory crosses a quarter of the entire BOM. The "compute" you thought you were buying is increasingly a memory system with a die attached.
The mitigation ladder: what operators actually do
When HBM is rationed, you cannot buy your way out โ so every serving trick that reduces KV bytes is, at bottom, an HBM procurement strategy. Ordered by leverage:
- Architecture (fixed at model choice). MLA compresses the cache 29.8ร against MHA and 3.73ร against GQA-8 (262,144 / 70,272 = 3.73): on the 192 GB variant that is 21.6 vs 5.8 concurrent 100k agents. Picking a GQA model in a memory-rationed market is a 3.7ร capacity self-penalty. The full derivation is in our KV-cache guide.
- KV quantization (2ร, config-only). FP8 KV halves the per-value byte cost: GQA-8 at 100k context drops from 26.2 GB to 13.1 GB per session โ the 192 GB GPU goes from 5.8 to 11.6 GQA agents. Int4 KV (e.g. Kimi's KIVI-style per-channel scheme) pushes further at accuracy cost.
- Prefix reuse (up to ~6.4ร on prefix-heavy fleets). RadixAttention-style sharing stores a shared system-prompt or few-shot KV once, not per session. A 2,000-token prompt is ~524 MB of GQA fp16 KV per concurrent user without sharing; with sharing, once.
- Tiering to host DRAM (unbounded, latency-priced). Offload cold KV pages over PCIe/CXL to the host's DDR5 โ which is itself the memory being crowded out by HBM wafer allocation, so note the irony. A 100k-context MLA session is 7.03 GB; a rack with terabytes of DDR5 holds thousands of parked agents, at the cost of re-read latency on activation. The full per-turn bytes-moved math is in our agent KV-tiering guide.
None of these add a single wafer. They are demand-side destruction of the only resource that is sold out โ which is precisely why they, and not new supply, set the operating economics until 2028.
Verdict
"AI demand is infinite" is doing no work in this analysis. Demand at what memory footprint? At MHA and 1M context, 0.00047 agents per GB means the world's entire HBM output serves a rounding error of sessions; at MLA and 100k context, 0.142 agents per GB means a single 192 GB GPU is a credible agent fleet. The shortage is real, but it is not a bolt-from-the-blue demand shock โ it is 3ร wafer-per-bit HBM economics meeting agentic workloads whose KV footprints are 10โ30ร chat workloads, priced through an allocation mechanism that cleared 2027 supply before the year began. New capacity is coming: SK hynix's M15X ramps toward ~50k wafers/month by mid-202712, Samsung's P4 phases in through 2026โ2027, and Micron's $9.3B-plus Idaho complex produces first DRAM wafers in H2 202712. That is long-lead physics โ cleanrooms, TSV tools, qualification cycles โ not press releases, and none of it relieves a 2026โ2027 serving plan. Plan in agents-per-GPU, not FLOPS. Buy architecture before you buy memory.
What to remember
- The shortage's transmission mechanism is the KV cache: 0.038โ0.14 agents per GB at agentic context means HBM capacity, not compute, is the product.
- ~3ร wafer area per bit (Micron, Hot Chips 2026) makes HBM supply structurally rationed independent of demand curve shape: capacity is chosen at wafer-start time, a year plus ahead.
- Rubin Ultra's reported 192/256 GB variants are a 33% capacity cut: 35.3 โ 21.6 MLA agents per GPU at 100k context โ peak FLOPS unchanged, residency gutted.
- $2,100 spot vs $300โ400 LTA is the allocation premium; HBM is now ~40โ47% of GPU manufacturing cost and ~25.7% of a VR200 rack BOM.
- The mitigation ladder โ MLA/GQA compression (3.73รโ8ร vs GQA/MHA), FP8 KV (2ร), prefix reuse (~6.4ร), host-DRAM tiering โ is the only lever that moves before 2028 capacity lands.
Footnotes
-
Micron at Hot Chips 2026, as reported by Tom's Hardware and Igor's Lab: HBM requires ~3ร the wafer area of DDR5 per bit, and the gap widens per generation. https://www.igorslab.de/en/micron-hbm-requires-three-times-wafer-area-ddr5-gap-widens โฉ
-
TrendForce, "Tight DRAM Supply Gives Suppliers Greater Pricing Power in HBM" โ HBM wafer input ~18%/22%/30% and HBM bit share ~8%/9%/13% of DRAM at end-2025/2026/2027; 2027 HBM4 contract prices expected to rise severalfold. https://www.dramexchange.com/WeeklyResearch/Post/2/12718.html โฉ โฉ2
-
Chosun Daily (Sept 8, 2026) and Sedaily (Sept 7, 2026), citing KB Securities: combined Samsung + SK hynix memory inventories below 10 days of supply. โฉ
-
Tom's Hardware: "Nvidia reportedly testing lower memory configs of Rubin Ultra as memory shortage bites back โ designs tested include as little as 192 GB and step back to HBM4," citing The Information; Nvidia response: "Our roadmap is intact." https://www.tomshardware.com/pc-components/gpus/nvidia-reportedly-testing-lower-memory-configs-of-rubin-ultra-as-memory-shortage-bites-back-designs-tested-include-as-little-as-192-gb-and-step-back-to-hbm4 โฉ
-
TrendForce (Aug 2026): Nvidia evaluating 8-Hi HBM4E, 12-Hi HBM4, and 8-Hi HBM4 alternatives for Rubin Ultra, citing 2027 DRAM tightness and 12-Hi HBM4E validation risk. โฉ
-
Intuition Labs (via Motley Fool / 247wallst.com), Sept 4, 2026: 36 GB HBM3E at ~$2,100 on the spot market. โฉ โฉ2
-
Seoul Economic Daily via Yahoo Finance: LTA pricing
2.87M won ($300โ400) per 36 GB HBM3E stack. โฉ -
Seoul Economic Daily via Yahoo Finance: ~70% of Samsung HBM capacity locked under LTAs with Microsoft, Nvidia, and Google, extending through 2031; SK hynix CEO expects shortage persistence through 2030. โฉ
-
TrendForce February 2026 forecast revision (via IBS analysis): conventional DRAM contract pricing +90โ95% QoQ in 1Q26 (upgraded from +55โ60%); Q2 2026 +58โ63%; Q3 2026 +13โ18%. โฉ
-
Widely-cited BOM teardown estimates (channel analysts): H100 ~$1,350 HBM of ~$3,320 manufacturing cost; B200 ~$2,900 HBM of ~$6,400. Direction corroborated by multiple independent BOM analyses; exact figures are estimates, not vendor-published. โฉ โฉ2
-
Morgan Stanley Research rack BOM estimates (Mayโ2026 supply-chain checks, corroborated across outlets): GB300 rack memory โ $374k of โ $3.99M; VR200 NVL72 rack memory โ $2.0M of โ $7.8M. โฉ โฉ2
-
SK hynix M15X (Cheongju): ~50k wafers/month by mid-2027, output committed largely to HBM; Micron Idaho ID1: first leading-edge DRAM wafers H2 2027; Samsung Pyeongtaek P4 lines phase in from 2026. Per Tom's Hardware, SK hynix investor materials, and Micron's US expansion disclosures. โฉ โฉ2