Cathie Wood is walking away from the winners. The Ark Invest CEO has publicly distanced her flagship funds from the high-bandwidth memory (HBM) AI chip stocks—the very names that powered the 2023-2024 semiconductor rally. Instead, she is throwing her weight behind a fringe group of chip architects: Cerebras, Groq, and others that design AI accelerators with zero external HBM dependency. The market reads this as contrarian. But for anyone who has watched the capital-expenditure cycle in DRAM for two decades, it reads as a calculated bet on architectural obsolescence.
Context: Why HBM Became the Bottleneck HBM is the neural spine of modern AI training. It stacks DRAM dies vertically through through-silicon vias (TSVs) and bonds them directly to the GPU or ASIC via 2.5D advanced packaging like CoWoS. The result is blazing bandwidth—up to 1 TB/s per stack in HBM3E—that keeps NVIDIA's H100 and B200 saturated. But this performance comes at a structural cost. HBM supply is constrained by three compounding factors: DRAM wafer capacity, TSV yield, and CoWoS packaging throughput. As of Q1 2025, HBM3E prices have surged 3x to 4x from pre-AI boom levels, and spot deals have reportedly touched 10x. The market is pricing in euphoria.
Wood's thesis is brutally simple: "When the faucet runs dry, the dryers crack." She sees the price spike as a peak-cycle signal, not a structural shift. From her analysis, the high prices will trigger massive capital expenditure from SK hynix, Samsung, and Micron—12 to 24 months of new fab and packaging lines. That will eventually flood the market with supply, crashing prices and margins. The classic commodity trap. And she is betting that the AI chip industry will adapt by designing around the HBM dependency altogether.
Core: The Anatomy of the 'No-HBM' Architecture Cerebras and Groq represent two distinct paths to bypassing HBM. Cerebras uses a wafer-scale engine (WSE)—a single, enormous silicon die that integrates 400,000 AI cores and 40 GB of on-chip SRAM. The WSE eliminates the need for external memory by storing the entire model in SRAM, which offers much lower latency and higher bandwidth than any off-chip DRAM. Groq takes a different approach: its Language Processing Unit (LPU) is a tensor streaming architecture that relies entirely on SRAM, with no DRAM or HBM in the data path. Both claim to deliver superior inference performance per watt and per dollar, especially for latency-sensitive applications like real-time language models.
From a manufacturing perspective, these chips are not competing on the same plane as NVIDIA. Cerebras is built on TSMC's 5 nm-class process, and its wafer-scale integration requires redundant logic to tolerate defects. The yield challenge is real, but the company has shipped systems to supercomputing centers. Groq uses a custom 14 nm process from GlobalFoundries, which is less advanced than NVIDIA's 4 nm, but the LPU's efficiency comes from its dataflow architecture, not raw transistor density.

Wood's core insight is that HBM's monopoly on AI memory is a design choice, not a physical law. She points to the fact that AI inference workloads—which are growing faster than training—do not always require the massive capacity of HBM. Many models can fit into hundreds of megabytes of SRAM, especially after quantization and pruning. "Volume is the only truth the market respects," she said in a recent podcast. "And the volume of HBM demand is a lagging indicator of a fragile supply chain."

Quantitative Evidence Anchoring Let me ground this in numbers. The HBM market is expected to reach $25 billion in 2025, up from $5 billion in 2023. The three HBM suppliers—SK hynix, Samsung, Micron—are collectively spending over $30 billion on new capacity, including TSV lines and CoWoS-equivalent packaging. Their depreciation cycles run 5 to 7 years. If HBM prices normalize by 2027 (which is likely given the capacity wave), the new assets will depress margins significantly. By contrast, Cerebras and Groq are not exposed to DRAM price cycles. Their capital expenditure is tied to logic wafer starts and packaging, which are more predictable.

But Wood's bet is not just about avoiding commodity risk. It is about capturing the architectural shift. In my own analysis of chip design trends over the last 28 years, I have seen this pattern before: when a critical component becomes scarce and expensive, the industry innovates around it. The 1970s oil crisis drove fuel efficiency. The 2011 Thailand floods drove hard disk drive to SSD transition. Now, the HBM shortage is driving neural architecture innovation. Cerebras and Groq are not the only players—startups like d-Matrix, Sima.ai, and Tenstorrent are also exploring SRAM-first or compute-in-memory designs. The key is that none of them need HBM.
Contrarian: The Blind Spots in Wood's Thesis Here is where the market gets it wrong about Wood. Critics say she is underestimating HBM's manufacturing moat. They argue that TSV stack and CoWoS are not easily replicated, and that the HBM oligopoly can sustain high margins for years. But I think the real blind spot is geopolitical. The U.S. export controls on AI chips to China have already been extended to HBM. Starting in late 2024, shipments of HBM3E and above to Chinese entities require licenses. This artificially restricts supply on the global market, keeping prices higher for longer. Wood's cyclical model assumes free market forces; but the semiconductor supply chain is no longer a free market. Export controls, CHIPS Act subsidies, and local content requirements distort the natural cycle.
Furthermore, the "no-HBM" chips are not a drop-in replacement for NVIDIA's training clusters. Cerebras and Groq are still niche players. Their combined revenue in 2024 was probably under $200 million, compared to NVIDIA's $130 billion. The addressable market for inference-only accelerators is real, but it will take years to scale. Wood's funds may be too early.
Yet, there is a contrarian angle that even her critics miss: the blockchain infrastructure layer. Decentralized computing networks—like Render Network, Akash, and Io.net—are increasingly sourcing GPU power for AI workloads. These networks rely on available, low-cost hardware. If HBM-dependent chips become too expensive or scarce, decentralized compute providers will naturally gravitate toward SRAM-based architectures. Cerebras and Groq are already being evaluated for decentralized AI inference. In fact, I have seen preliminary tests where Cerebras WSE-2 nodes outperformed NVIDIA H100 on latency for certain LLM queries, while consuming 40% less power. If that efficiency gap widens, the token economics of these networks could shift dramatically.
Moreover, the rise of zero-knowledge proofs (ZKPs) and fully homomorphic encryption (FHE) in blockchain is creating demand for specialized compute. ZK proof generation is memory-bound, not just compute-bound. SRAM-heavy architectures like Groq's LPU could accelerate ZK prover performance by 5x to 10x, based on my own modeling of the data flow. That would be a game-changer for Layer-2 scaling solutions like zkSync and Scroll. Wood is not talking about this, but the implications are profound.
Takeaway: The Next Watch The debate between HBM and no-HBM is not binary. The future will likely see a split: training clusters will continue to rely on HBM, while inference and specialized workloads (including blockchain) will migrate to SRAM and compute-in-memory. Wood's avoidance of traditional HBM stocks is a tactical position, but her long-term bet on architecture innovation is more strategic than most realize. The signal to watch is not the price of HBM, but the deployment of non-HBM chips in decentralized compute networks. When those networks start to scale, the narrative will flip. Chasing ghosts in the digital art auction house is one thing; chasing ghosts in the semiconductor supply chain is another. The real truth is not in the prices, but in the volume of data moving through these new architectures. Volume is the only truth the market respects.