On June 4, SemiAnalysis reported that NVIDIA is lowering the standard configuration of SOCAMM, the memory module attached to the CPU side of Vera Rubin NVL72.
The standard build, which had been expected to center on 192GB modules, shifts to 96GB modules, and CPU-side memory per rack drops from roughly 54TB to 28TB, cut in half. The next day, memory and storage stocks fell across the board, Micron and SanDisk among them.
The questions followed immediately.
If CPU orchestration is supposed to matter more, why cut CPU-side DRAM in half first?
HBM4 capacity is not rising much over the previous generation either, so isn’t this a sign that memory demand is weakening?
The last three pieces each took on one slice of this question.
The starting point of the first piece was that the bottleneck in AI inference is moving outside the GPU, toward the CPU and system memory.
The second piece made the point that a bottleneck forming does not mean the company sitting on that bottleneck immediately captures the profit.
The third piece asked how much of AI memory demand actually has to stay inside HBM.
Vera’s DRAM cut is the event where those three slices meet at a single point.
The migration of the bottleneck, the distribution of profit, and the demand that has to remain in HBM are all put to the test at once, inside the memory configuration of a single rack.
This piece starts with how far that 28TB number should be trusted. From there it follows whether the half of demand that appears to have vanished has actually vanished, whether HBM, SOCAMM, and storage were ever things that belonged together for the same reason, and what the reduced DRAM is really pointing at.
What the market priced in and what the SemiAnalysis report actually said were not the same thing.
Disclaimer
This piece reflects a personal opinion based on public information and does not recommend buying or selling any specific security. Responsibility for investment decisions and their outcomes rests with the investor.
1. The First Misreading the DRAM Cut Created
NVIDIA has announced nothing.
The official reference spec still lists up to 54TB, and what SemiAnalysis reported is a supply chain account that the actual shipping standard build is dropping toward 96GB. Eight SOCAMM modules attach to each Vera CPU, and a rack holds 36 Vera CPUs, so 192GB modules line up with the official spec’s 54TB, while 96GB modules come to roughly 28TB.
The background the report gave is supply and cost.
With the supply of LPDDR5X chips going into SOCAMM modules not keeping up enough to fill the full standard build, the move is to lower the standard configuration to 96GB and get racks installed first rather than delaying rack shipments over memory, leaving the high-capacity build as a custom option.
By the report’s figures, this adjustment lowers the cost of a single rack from about $7.6M to $6.8M, and TCO per GPU comes down with it. So 54TB is the number that says how far it can go, and 28TB is the report of how it actually gets installed within that constraint.
Still, the reason investors reacted to this number is understandable.
All along we have been saying that CPU orchestration matters more in AI inference. If the CPU feeds the GPU cluster, splits up requests, and moves state, meaning the context and intermediate results a model has to hold while it works, between memory and storage, then why is CPU-side DRAM shrinking?
And in the market this doubt got translated into a simpler sentence.
If DRAM dropped in Vera, hasn’t memory demand weakened?
Take that question as is and the answer scatters into product-by-product verdicts.
You end up arguing whether to read the DRAM cut as bearish, the HBM4 hold as bullish, the SOCAMM mix one way or the other.
Go down that road and the question fragments at every product name, and the force that actually matters disappears.
The question has to change.
Of the AI inference state, which part must stay inside HBM4, and which part can move outside it?
2. The Fastest Lane Did Not Weaken
What shrank is the CPU-side memory, not HBM4.
GB300 NVL72 put forward 20.7TB of HBM3e and 576TB/s of GPU memory bandwidth.
Vera Rubin NVL72 puts forward 20.7TB of HBM4 and 1,580TB/s of GPU memory bandwidth.
Capacity did not rise much, but bandwidth gets far stronger.
HBM4’s change is in the speed of the corridor rather than the size of the warehouse.
In inference, not all data carries the same temperature.
The hot state the GPU has to read right away is more sensitive to bandwidth and latency. The hot KV that gets referenced immediately for the next token among the KV cache where the model stores the context it has read so far, the intermediate state that has to move back and forth quickly, the data that has to keep flowing so the GPU does not stall, all of that belongs here.
As long as this kind of state remains, the investment logic for HBM4 is hard to judge by simple byte growth alone. The fact that capacity does not rise much can be disappointing. But the fact that bandwidth climbs by nearly three times also means NVIDIA is not loosening this lane.
The supply side’s remarks line up with this direction too.
Micron referenced HBM4 high-volume production and SOCAMM2 volume shipment for Vera Rubin together.

Samsung also referenced HBM4 and SOCAMM2 productization for the Vera Rubin generation in its first quarter 2026 results and its GTC 2026 announcement.

This does not prove any particular rack deployment mix. But it reads clearly enough as a supply-side signal that vendors are preparing the HBM4 hot lane.
In other words, the debate over Vera’s DRAM cut is not evidence that the HBM4 thesis has disappeared.
You have to first acknowledge that HBM4 remains the faster, more expensive hot lane.
Only then does the second question open up.
If HBM4 is that important, why does the memory outside HBM become more important?
3. A Strong Hot Lane Creates a Buyer Problem
That HBM4 is fast and strong is not purely good news for the platform owner. It is an expensive lane with tight supply, so you cannot put everything on it. That much more, you have to choose harder what to keep on HBM4 and what to push outside it.
That choice starts from state. State is the working data a model has to hold while it reasons, which is context, KV cache, and intermediate results. Not all state carries the same weight. Hot state that has to be reached again quickly with every token will idle the GPU without a fast lane like HBM4. State that gets touched occasionally, or holds up even when a bit slower, can be moved down to a cheaper, slower tier without hurting performance much.
Deciding this placement is the orchestration the CPU handles. But saying the CPU matters more does not mean CPU-side DRAM has to get bigger. Its importance lies in the power to control what goes where, rather than in capacity. Holding the same memory, getting this placement right saves expensive HBM while keeping the GPU busy, and getting it wrong leaves the GPU idle no matter how much memory you add.
Think it through again from the platform owner’s position with this view, and Vera’s DRAM cut is not a strange event. It is the platform owner concentrating the expensive HBM4 lane on hot state and testing how far the less hot state can be sent outside.
So does memory demand drop by however much got cut?
The SOCAMM bits that ship out are not set by the capacity of one system alone.
SOCAMM bit demand = module capacity x module per system x system shipped
It is the product of these three: module capacity, modules per system, systems shipped.
96GB lowers the first term, halving the per-system content of one system. But if the lowered capacity loosens the supply constraint and more systems get installed, the last term rises. The number of systems can make up for the amount cut per system.
So reading 96GB straight as a memory demand collapse is premature, and so is reading it as a trivial launch optimization.
Which one it is can only be known by watching how many systems actually get installed in what configuration.
Get this far and the three original questions converge into one.
DRAM dropped, so has memory demand turned down?
Not knowable yet.
If the CPU matters, why cut DRAM?
If the CPU’s job is to decide where state goes rather than to grow the local DRAM total, it is not a contradiction.
If HBM4 gets stronger, where does the state outside it go?
That question is exactly the captive boundary.
From here on we follow state, not products. State held captive inside HBM4, state absorbed within the DRAM family, state that can fall to a slower tier. Here the DRAM family means the DRAM line outside HBM, namely CPU-side LPDDR and SOCAMM.
Depending on where the boundary between these three is drawn, the profit pool gets trapped inside HBM, absorbed into SOCAMM, or leaks out to storage. So even the same memory long buys completely different things depending on which box you stand in. The exposure differs by box too much to bundle under the one word memory, and among them are names with almost no exposure to this cut, and names for which the cut actually widens the field.
Where that boundary sits now, who can move it, and where that difference leaves opportunity, that is the real content of this event.







