Your AI Stack Is Hitting a Wall and Most Teams Aren’t Ready

The memory wall isn’t coming. It’s here. And if you’re running GPU workloads at any meaningful scale, it’s already costing you.
Recently, Understanding AI reporter Kai Williams sat down with Val Bercovici from WEKA to cut through the noise: The constraints facing AI aren’t theoretical, the memory crisis is active and growing, and the window to stay ahead of the curve is narrowing quickly.
The “IO blender” problem and what it means for AI teams
Most infrastructure conversations fixate on GPU counts and network bandwidth. This is the wrong place to look. One of the most punishing — and underappreciated — obstacles in AI has always lived at the storage layer. When training models, engineers aren’t dealing with a handful of large, uniform files. They’re wrangling billions of files with wildly different sizes scattered across systems, generating a chaotic and unpredictable I/O load that breaks conventional architectures.
As Val explained: “When you’re training models, it’s not off of just a bunch of files; it’s billions of files, and they’re not the same size. They’re not just large, they’re not just small, and they’re scattered everywhere. So there are metadata requirements. We call this the ‘IO blender.’”
Legacy file systems like Lustre required full-time teams just to keep pace. The operational lesson for today’s infrastructure leaders deploying much more complex AI systems is unambiguous: Storage architecture is not a secondary concern, and teams must take it seriously so it can be a force multiplier for AI instead of a drag on every dollar of compute you’ve invested.
Why AI inference has completely different infrastructure requirements
Post-ChatGPT, the industry’s center of gravity shifted from training to inference, and the two disciplines demand fundamentally different approaches. Training resembles building a supercomputer: centralized, tightly coupled, batch-oriented. Inference is a different animal entirely.
“Inference is very, very different,” Val said. “It’s much more stateless if you do it right. It's much more decentralized than centralized.”
This decentralized nature has also fractured latency requirements into distinct tiers that can no longer be served by a one-size-fits-all stack. Real-time voice agents demand fixed, predictable latency. Chat and research sessions tolerate a middle tier. Agentic swarms — now routinely running for hours and even days — need throughput optimization to ensure compounding turn delays don’t erode what should be a competitive advantage into an operational liability. Infrastructure teams still treating inference as a monolithic workload are building for a world that no longer exists.
The memory wall: AI infrastructure’s most urgent (and active) crisis
Here is the conversation’s sharpest warning: The industry has already hit the memory wall. This isn’t a forecast; it’s today’s reality. Model weights for large language models (LLMs) combined with exploding KV cache demands from concurrent agent sessions have already outpaced available memory capacity, and the problem accelerates as agentic workloads multiply.
“The demands of model weights for large language models, coupled with large KV caches per user per agent session in a swarm, have literally hit a wall right now,” Val said. “Instead of prefilling once logically, you’re prefilling thousands and millions of times for agent swarms redundantly.”
This redundant recomputation is burning GPU cycles (and GPU budget) on work that better memory architecture would eliminate entirely. If you are running agentic pipelines at scale, this isn’t a performance nuisance. It’s a direct, measurable tax on every workload you’re running.
Supply constraints mean software efficiency is now an essential survival skill
Compounding the memory problem, the hardware supply chain offers no near-term relief. New fabrication capacity for GPUs, DRAM, HBM, and NAND flash requires multi-year, multi-billion-dollar investments that increased market demand simply cannot accelerate. Val noted that some speculative panic-buying inventory is cycling back into the market, creating unpredictable ups and downs in price that signals jaggedness rather than a straight line up. But he was clear that no one should mistake short-term fluctuations for a fundamental fix.
The only lever you can actually control is efficiency. “The solution isn’t software. The solution is in being more efficient with the very scarce resources you have,” Val said. “We’re running AI factories before the assembly line, before the Model-T moment.”
This reframe matters. Performance density — output per watt, per GPU, per dollar — is no longer a hyperscaler concern. It is now a competitive variable for any organization running inference at meaningful scale, and the gap between teams that have internalized this and those that haven’t is widening by the quarter.
The agent inflection point that changed the calculus permanently
Val laid out a timeline that deserves to be taken seriously. At GTC 2025, inference was underemphasized and agents were absent from production. By May 2025, just two months later, tools like Claude Code signaled what real agentic capability could look like. By December 2025, what Jensen Huang called an inflection point had arrived, and token demand stopped having a visible ceiling.
“All of a sudden, there’s no end to the rise and growth in token demand,” he said.
The uncomfortable truth is that most teams won’t act until the cost is undeniable. GPU cycles burned on redundant KV cache recomputation. Agent pipelines that compound latency into competitive disadvantage. Inference tiers that weren’t designed for the workloads now running on them.
The teams pulling ahead in 2026 aren’t waiting on better hardware. They’re auditing their memory architecture now, separating their inference tiers and treating software efficiency as a first-order discipline instead of an optimization they’ll get to later.
The memory wall doesn’t care about your roadmap. The question isn’t whether it’s affecting your stack. It’s whether you’ll find out on your terms or your infrastructure’s.
Watch the full GTC 2026 conversation with Val and Kai on YouTube.
What's Next
Scale Production AI Faster with NeuralMesh
Your models aren't slow. Your data is. Fix AI bottlenecks with high-throughput infrastructure.


