# How WEKA’s Augmented Memory Grid™ brings petabytes of persistent storage for KV Cache

**Published:** May 15, 2025

![How WEKA’s Augmented Memory Grid™ brings petabytes of persistent storage for KV Cache](https://cdn.sanity.io/images/ult5g8gw/production/8cad9f85af37539c414835d5cf295e4e52db900e-960x540.jpg)

**Watch:** https://fast.wistia.net/embed/iframe/xxpnzb6wcm

Learn how Augmented Memory Grid radically improves the economics and performance of AI inference.

## Commentary

Ever wondered how large language models (LLMs) handle your questions behind the scenes? In this demo, Callan Fox from WEKA walks you through a real-world AI inference scenario: uploading “The Martian” to an LLM to fact-check scenes from the movie.

## Related Videos

- [Why AI Inference Margins Are Suddenly Getting Better](/video/why-ai-inference-margins-are-suddenly-getting-better)
- [Inference Is Eating Memory, and Tokens Now Run AI Economics](/video/inference-is-eating-memory-and-tokens-now-run-ai-economics)
- [Build for What's Next in AI with NeuralMesh](/video/build-for-what-s-next-in-ai-with-neuralmesh)
