I watched fortunes bloom and wither in real-time. Today, I am watching a different kind of bloom—a potential withering of the very architecture that powers the AI gold rush. The market is screaming about HBM prices, about SK Hynix and Micron's soaring revenues. But the code doesn't lie, and the architecture whispers a different story. Cathie Wood’s recent pivot away from high-bandwidth memory (HBM) dependent AI chip stocks, and her endorsement of ‘HBM-less’ architectures like Cerebras and Groq, isn't just a contrarian trade. It's a signal. A signal that the next phase of the AI chip war won't be fought with more memory stacks, but with smarter, more integrated silicon. The question isn't whether HBM is overpriced. The question is whether the architecture it enables is becoming a liability.
The Context: The HBM Mirage
To understand the bet, you must first understand the battlefield. HBM isn't just a memory chip; it's a system. It's a stack of DRAM dies, connected through thousands of Through-Silicon Vias (TSVs), bundled with a logic die, and then packaged onto a GPU or ASIC using advanced 2.5D/3D packaging like CoWoS. This is the spine of the modern AI training cluster. For the last three years, this spine has been getting stronger—and more expensive. The HBM market is a classic oligopoly: SK Hynix, Samsung, and Micron hold the keys. They are the gatekeepers of AI memory bandwidth. And they have been raising the toll.
Wood’s argument is deceptively simple: HBM is a commodity. Its price has surged 3x, 4x, even 10x in some reports. This isn't a sign of structural demand, she argues. It's a sign of a cyclical peak. The price signal is a warning. High prices will trigger a wave of capital expenditure, which will flood the market with supply in 12-24 months, collapsing prices. More importantly, high prices will incentivize chip architects to find alternatives. This is the core of her thesis: the very success of HBM will sow the seeds of its own destruction. The industry will design around it.
But is that true? In my experience building real-time trading models, I've learned that the market often confuses price with value. The price of HBM is high because the demand for AI training is structurally real and immense. But the value of HBM is being challenged. The architecture of an AI accelerator is a delicate balance of compute, memory bandwidth, and memory capacity. HBM provides massive bandwidth, but at a cost. It's a separate chip, requiring a complex and expensive packaging process. It introduces latency. It consumes power. It creates a supply chain chokepoint. This is the vulnerability that Wood is betting on.
The Core: The Architecture of Defiance
Let's get into the silicon. The two primary challengers to the HBM-status quo are Cerebras and Groq. They don't just use different chips; they use different physics. Cerebras famously builds a Wafer-Scale Engine (WSE). Instead of cutting a wafer into dozens of individual dies, they use the entire wafer as one gigantic chip. The current WSE-3 is built on a 5nm process, packing 4 trillion transistors. It's a monstrous chip. The genius, however, isn't just the size. It's the memory. The WSE integrates 44 GB of on-chip SRAM, distributed across its 900,000 cores. This SRAM is the secret weapon. It's incredibly fast, with a bandwidth of 21 PB/s—orders of magnitude faster than any external HBM connection. The latency is measured in nanoseconds, not microseconds. By keeping the data on the chip, Cerebras eliminates the need for the HBM bridge entirely.
Groq takes a different but equally radical path. Their Language Processing Unit (LPU) is an architecture built for deterministic performance. It uses a tensor streaming processor (TSP) architecture, where the compute and memory are tightly coupled. Again, the key is on-chip SRAM. Groq’s LPU uses a massive SRAM pool to feed its compute units, avoiding the need for a complex HBM memory hierarchy. The result is an architecture that excels at inference, especially for large language models, with extremely low latency and predictable performance.
Both of these architectures are answers to a fundamental question: is the Von Neumann bottleneck—the separation of memory and compute—worth it? For decades, we've accepted it. But HBM is a very expensive, very complex, and very power-hungry way to bridge that gap. Cerebras and Groq are saying: what if we just don't cross the bridge? What if we build the memory into the compute?
This is the core of my analysis. Based on my audit experience of DeFi protocols, I've seen how a single point of failure can cripple a system. HBM is a single point of failure for the NVIDIA ecosystem. It's a supply chain chokepoint. The architecture of defiance is not just about performance; it's about resilience. It's about building a system that is less dependent on a complex, concentrated, and cyclical supply chain.

The market is currently pricing AI chips based on the assumption that HBM is the only path. Cerebras and Groq are proving that's a false assumption. The data shows that for specific workloads—especially inference—the SRAM-based architectures can match or even exceed the performance of HBM-based systems while consuming less power and having a simpler supply chain. This is not a theoretical argument. Cerebras has deployed systems with the likes of Mayo Clinic, GlaxoSmithKline, and the US Department of Energy. Groq is powering inference for a growing number of enterprise customers. The evidence is there, but it's being ignored by a market fixated on the NVIDIA/HBM narrative.
The Contrarian Angle: The Hidden Cost of the Boom
Here is the unreported angle. The market is cheering HBM price increases as a sign of strength. But it should be asking: what is this price doing to the ecosystem? The price of HBM is not just a line item on a BOM. It is a signal that is reshaping the entire AI chip landscape. The market is missing the fact that high HBM prices are acting as a massive, indirect subsidy for the ‘HBM-less’ architectures. Every dollar that NVIDIA pays for HBM is a dollar that could have been spent on developing a better on-chip memory solution. The high price is creating a massive incentive for innovation in the non-HBM space.
Speed is survival, but empathy is the signal. Let me show you the empathy here. The real victims of this price surge are not the hyperscalers. They can afford it. The victims are the mid-tier AI companies, the startups, the researchers. They are the ones who are being priced out of the AI race. They are the ones who are desperate for a more affordable, more accessible AI compute solution. The Cerebras and Groq architectures are not just faster; they are potentially cheaper. By eliminating the HBM and the complex CoWoS packaging, they can offer a lower total cost of ownership. This is the human story behind the architecture. It's not just about the technology; it's about access. It's about making AI compute more democratic.
Furthermore, the analysis of the supply chain reveals a deeper vulnerability. The real bottleneck is not the DRAM itself, but the advanced packaging and TSV capacity. The HBM supply chain is a triple-chokepoint: DRAM supply, TSV capacity, and CoWoS capacity. A disruption in any one of these can cripple the entire AI chip supply chain. The HBM-less architectures bypass this entire complex web. They are not just a different chip; they are a different, simpler, and more resilient supply chain. The market is currently pricing in the risk of a HBM shortage, but it is not pricing in the opportunity of a HBM-less architecture that is immune to it.

The Takeaway: The Bet on the Bifurcation
Cathie Wood is not betting that HBM will crash. She is betting that the AI chip market will bifurcate. One branch will continue to be the HBM-dependent, NVIDIA-dominated training beast. The other branch will be an architecture-defined, HBM-agnostic, inference-first world. The training market is huge, but it is also a known quantity. The inference market is the next frontier. And it is in the inference market where the ‘HBM-less’ architectures have their strongest advantage. Low latency, high throughput, deterministic performance, and lower power consumption are the critical requirements for real-time AI inference. These are the exact requirements that Cerebras and Groq are designed to meet.
The code didn't change. The market did. The architecture of the AI chip is not a settled science. It is a dynamic, evolving field. The next major breakthrough might not come from a faster transistor or a denser memory stack. It might come from a smarter way to organize the pieces we already have. The bet on Cerebras and Groq is a bet on architectural innovation. It's a bet that the future of AI compute will be defined not by the size of your memory stack, but by the elegance of your design. Stability isn't a line of code. It's a careful architecture. The architecture of the future is one that is less dependent on the past. The signal is clear. The HBM era is not ending, but a new era is beginning. The architects are at the gate. The question is: are we listening?
I watched fortunes bloom and wither in real-time. The next fortune will be built by those who understand that architecture is destiny. The code is the law, and I am its restless guardian.
