Skip to content
← Back to Active Thesis Feed
🟢 Confirmed Aug 27, 2026

Meta Is Running Memory It Was About To Throw Away

Hyperscalers are pulling DDR4 modules out of servers headed for scrap and running them behind CXL controllers to keep them in production. Meta is doing it across millions of servers, cutting server counts by up to 25% on some inference workloads. That is not a cost-optimization program. It is what a business does when it cannot buy the part. Memory has gone from roughly 8% of hyperscaler capital spending in 2023 and 2024 to about 30% this year, conventional DRAM contract prices rose 90% to 95% in a single quarter, and NVIDIA's CFO used the word "extreme" on the earnings call and said prices are heading higher into next year. The AI buildout has a bottleneck, and it is not compute.

NVIDIA CFO Colette Kress, on the August 26 earnings call: "We are experiencing extreme pricing conditions in memory. The magnitude of the price increase has exceeded our prior expectations and are headed even higher into next year." She separately volunteered that NVIDIA has "longstanding, deep relationships with all three major memory suppliers" — a reassurance no analyst asked for. In the same release, NVIDIA's total commitments rose to $279 billion from $119 billion quarter over quarter. The company attributed the increase primarily to memory procurement. That is $160 billion of incremental contractual obligation in ninety days to secure one input. Gross margin was 75% in the quarter and guided to 73.5%–74.5% for Q3, while NVIDIA has separately told its largest customers that AI server prices will rise more than 15% in many cases, citing memory costs. On the buyer side, Marvell is pushing DDR4 recycling using CXL technology. Hyperscalers are placing DDR4 modules behind CXL controllers to extend the useful life of memory that would otherwise be scrapped when servers move to newer DDR5-only platforms. Memory will account for roughly 30% of hyperscaler capex this year, up from about 8% in 2023 and 2024. Conventional DRAM contract prices rose 90% to 95% in a single quarter. Meta is already running recycled DDR4 behind CXL across millions of servers, cutting server counts by up to 25% for some inference workloads.
View source ↗ 2026-08-26
The framework has tracked memory tightness for two months through a series of increasingly indirect signals, each of which had a plausible innocent explanation. Micron said it was filling less than half of what data center customers requested — but requests are free, and buyers pad orders to game allocation. Meritz reported four buyers asking for 95 to 100 billion gigabits of 2028 HBM against a total market forecast near 55 billion — same objection, twice as loud. SK hynix began refunding cash instead of shipping replacement SSDs, which is suggestive but is a company policy rather than an industry condition. Samsung raised foundry prices 10% to 15%, which only works if TSMC is booked out. None of those individually settles it. This does, because it is the first evidence that carries a price and a penalty rather than an intention. A purchase commitment is a signed contract. NVIDIA did not request $160 billion of memory. It obligated itself to buy it, in a single quarter, and told investors the price is worse than it expected and getting worse still. Companies do not lock in contractual exposure of that size for inputs they can buy on demand. The recycling is the other half and it is the more revealing one. Meta pulling DDR4 out of decommissioned fleets and running it behind CXL controllers means the alternative — buying new DDR5 — is either unavailable or uneconomic at current prices. Cutting server counts by 25% by putting more memory behind fewer machines is a workaround for a memory constraint, not a compute one. When a company with Meta's balance sheet chooses scrap over purchase orders, the shortage has stopped being a pricing question. Which reframes the Rubin Ultra roadmap change. NVIDIA's flagship successor was previewed with 1TB of HBM and will ship with 192GB — less than base Rubin's 288GB — after the die count was halved, the stack height cut from 16-hi to 8-hi, and the generation stepped back from HBM4E to HBM4. The available readings were that HBM supply forced the redesign, or that a 3,600W thermal envelope proved undeliverable and memory content followed the smaller package. A company committing $160 billion incrementally to secure memory while calling conditions extreme is not a company with a free hand on content. They cut content per unit and pre-bought supply. Both are rationing. The number that should worry anyone underwriting the buildout is the capex share. Memory at 30% of hyperscaler capital spending against 8% two years ago means nearly a third of the AI budget is now going to a component whose price rose 90%-plus in a quarter. Against $1.52 trillion of purchase commitments and $904 billion of leases not yet started across Alphabet, Amazon, Meta and Microsoft, an input cost moving like that does not reduce the obligation. It reduces what the obligation buys.
  • The clean falsifier: DRAM contract prices flattening or declining in the next quarterly pricing round. The shortage thesis has been built on rate of change, and the rate of change is the thing that breaks first.
  • Whether Micron, SK hynix or Samsung announce accelerated capacity. Micron's meaningful DRAM additions are a 2028 event — Tongluo greenfield with EUV, 1D2, Hiroshima Fab 15 — with 2027 limited to roughly 80k wafers per month at 1D1 and 40k at Tongluo brownfield. If that timeline pulls forward, relief arrives earlier than the commitments assume.
  • Whether other hyperscalers disclose CXL recycling programs. Meta doing it is one company's engineering choice. Three doing it is an industry condition, and it would also mean DDR5 unit demand is weaker than the price implies.
  • Watch NVIDIA's commitments line next quarter. If $279 billion grows again, the procurement is not finished. If it flattens, they believe they have covered their supply.
  • The counterweight worth holding honestly: recycling and CXL are genuine efficiency technologies with a rationale independent of scarcity, and 25% fewer servers for the same inference work is a real cost saving any operator would pursue in a normal market. The scarcity reading rests on the timing and on the price data, not on the practice itself.
  • And the second-order risk nobody is pricing: if a quarter of inference server demand can be met by re-racking existing memory, that is unit demand not placed with the OEMs. The efficiency that solves the shortage also shrinks the order book.
⚡ The AI buildout's binding physical constraint is now documented as memory rather than compute, with a contractual number attached, and the input cost is inflating inside commitments that cannot be renegotiated.