# WEKA Videos

Videos from WEKA on data infrastructure, AI/ML storage, and high-performance computing.

## [Inference Is Eating Memory, and Tokens Now Run AI Economics](/video/inference-is-eating-memory-and-tokens-now-run-ai-economics)

**Published:** July 29, 2026

Expanding agent contexts stall GPU inference. Hear how Augmented Memory Grid offloads KV cache to NVMe for 90% savings.

## [Build for What's Next in AI with NeuralMesh](/video/build-for-what-s-next-in-ai-with-neuralmesh)

**Published:** July 21, 2026

AI needs a foundation that keeps up with it. NeuralMesh™ 6 runs training, inference, and agentic workloads on one adaptive platform. Built to scale without slowing you down.

## [WEKApod: Maximum Capacity and Performance Density](/video/wekapod-maximum-capacity-and-performance-density)

**Published:** July 21, 2026

The teams winning in AI aren't waiting for more space or power. They're getting more from what they already have. WEKApod™ packs it all into one rack. 


## [The Inference Economy Runs on WEKA](/video/the-inference-economy-runs-on-weka)

**Published:** July 21, 2026

Behind every stalled GPU is a cure, a breakthrough, a discovery waiting to happen. WEKA built the foundation to make sure AI doesn't have to wait for it.

## [What It Actually Costs to Run Intelligence at Scale](/video/what-it-actually-costs-to-run-intelligence-at-scale)

**Published:** July 20, 2026

Uncapped token costs stall enterprise AI scaling. Hear experts map how a 100x unit cost reduction drives a 10,000x demand surge.

## [Keeping GPUs Busy Is the Real AI Infrastructure KPI](/video/keeping-gpus-busy-is-the-real-ai-infrastructure-kpi)

**Published:** July 16, 2026

Idle GPUs stall AI workloads. Hear HTX's Chief Innovation Officer explain how optimized infrastructure shortens the data circuit to maximize GPU utilization.

## [Why Memory, Not Compute, Is Breaking the AI Cost Curve](/video/why-memory-not-compute-is-breaking-the-ai-cost-curve)

**Published:** July 14, 2026

GPU memory limits stall active agent swarms. Augmented Memory Grid delivers 10x more sessions without more GPUs.

## [Why the AI Memory Wall Depends on Storage, Not DRAM](/video/bypassing-the-ai-memory-wall-val-bercovici-on-the-generalists-podcast)

**Published:** July 7, 2026

KV cache evictions stall long-horizon AI agents. Augmented Memory Grid bypasses DRAM to storage and drives 10x more concurrent tokens.

## [Bypassing the Inference Memory Wall to Scale Agentic AI](/video/bypassing-the-inference-memory-wall-to-scale-agentic-ai)

**Published:** June 30, 2026

Dan Nishball, Director of Research at SemiAnalysis and Val Bercovici, WEKA Chief AI Officer, break down the token demand explosion, the memory wall slowing AI inference, and why cybersecurity agent swarms may be enterprise AI’s next frontier.

## [The Data Path Is the Mission](/video/the-data-path-is-the-mission)

**Published:** June 16, 2026

Legacy storage is the choke point holding federal AI back. See how WEKA NeuralMesh™ delivers mission-speed data for training, inference, and the tactical edge.


## [Storage Is the Spinal Cord of AI: Inside Singtel’s Sovereign AI Cloud](/video/storage-is-the-spinal-cord-of-ai-inside-singtel-s-sovereign-ai-cloud)

**Published:** June 10, 2026

Singtel's CTO explains why storage is the spinal cord of their sovereign AI cloud — and how WEKA powers GPU clusters from H100s to GB200s.


## [The Real Cost of AI: How Smart Companies Are Maximizing Token ROI](/video/the-real-cost-of-ai-how-smart-companies-are-maximizing-token-roi)

**Published:** May 19, 2026

WEKA Chief AI Officer Val Bercovici joins investors and cloud leaders at HumanX to debate AI token costs, procurement strategy, infrastructure moats, and what "return on intelligence" really means.


## [How to Get Signalmaxxing Out of Tokenmaxxing](/video/how-to-get-signalmaxxing-out-of-tokenmaxxing)

**Published:** April 29, 2026

Optimize your AI tokenomics. Scale signalmaxxing out of tokenmaxxing without tripling token per watts or risking hardware-driven model latency.

## [Your GPUs Are Waiting. The IO Blender and Memory Wall Explain Why.](/video/your-gpus-are-waiting-the-io-blender-and-memory-wall-explain-why)

**Published:** April 21, 2026

Stop wasting expensive GPU resources. See how the memory wall and IO blender trigger massive I/O wait time and a neural network training bottleneck.

## [The Future of Frontier Models And What They Will (And Won’t) Do Next](/video/the-future-of-frontier-models-and-what-they-will-and-won-t-do-next)

**Published:** March 26, 2026

Stop letting storage bottlenecks stall your AI factory. Optimize AI infrastructure investment and GPU utilization to scale next-gen frontier models.

## [Why AI Storage Will Define the Future of Inference at Scale](/video/why-ai-storage-will-define-the-future-of-inference-at-scale)

**Published:** March 20, 2026

Stop wasting GPU spend. Discover how optimizing TCO for AI storage eliminates memory bottlenecks and dramatically improves enterprise inference ROI.

## [Building Superintelligence with Zyphra’s Chief AI Strategy Officer](/video/building-superintelligence-with-zyphra-s-chief-ai-strategy-officer)

**Published:** March 10, 2026

Eliminate GPU memory limits. Zyphra's Erik Norden shares how optimizing KV cache hit rate prevents latency bottlenecks in agentic AI superintelligence.

## [LinkedIn’s AI Infrastructure Secrets for 1.2B Users](/video/linkedin-s-ai-infrastructure-secrets-for-1-2b-users)

**Published:** March 3, 2026

Eliminate idle GPU bottlenecks to slash costs. Discover how LinkedIn optimizes GPU utilization and storage to power AI models for 1.2B users efficiently.

## [Why Inference Will Drive AI Infrastructure in 2026 with Crusoe](/video/why-inference-will-drive-ai-infrastructure-in-2026-with-crusoe)

**Published:** February 24, 2026

Solve GPU memory bottlenecks to boost Inference ROI. Crusoe and WEKA discuss why AI infrastructure investment must evolve to optimize TTFT and latency.

## [Driving AI Innovation While Scaling Capacity with Meta](/video/driving-ai-innovation-while-scaling-capacity-with-meta)

**Published:** February 17, 2026

Solve GPU underutilization. Optimize GPU utilization and achieve GPU cost reduction to keep hardware procurement cycles from stalling AI capacity scaling.

## [AI Token Economics and the Real Cost of Running AI Models](/video/ai-token-economics-and-the-real-cost-of-running-ai-models)

**Published:** February 10, 2026

Optimize your AI infrastructure investment. Overcome GPU memory constraints to lower tokenomics cost and accelerate model scaling.

## [GPU Capacity Planning and Compute Market Dynamics](/video/gpu-capacity-planning-and-compute-market-dynamics)

**Published:** February 5, 2026

Avoid GPU resource limits. Align your AI infrastructure investment with compute dynamics to maximize GPU utilization and prevent system latency.

## [Optimizing AI Inference Economics | A Conversation with SemiAnalysis](/video/optimizing-ai-inference-economics-or-a-conversation-with-semianalysis)

**Published:** February 2, 2026

Boost GPU utilization and slash scale-out storage costs. Solve memory bottlenecks to maximize inference ROI and scale enterprise AI models efficiently.

## [AI Infrastructure Management and ROI Measurement](/video/ai-infrastructure-management-and-roi-measurement)

**Published:** January 28, 2026

Stop wasting GPU utilization on slow storage. Maximize your AI infrastructure investment to prevent critical model latency and secure enterprise ROI.

## [The Agentic AI Infrastructure Playbook](/video/the-agentic-ai-infrastructure-playbook)

**Published:** January 26, 2026

Optimize your AI infrastructure investment. Solve GPU memory limits and improve Time to First Token (TTFT) to unlock maximum AI inference efficiency.

## [Intelligence Unlocked: The Data Foundation Fueling AI Factories](/video/intelligence-unlocked-the-data-foundation-fueling-ai-factories)

**Published:** January 21, 2026

Accelerate your AI factory. See how optimized AI infrastructure investment eliminates latency bottlenecks to fuel faster enterprise token generation.

## [Creating Scalable AI Factories Built for Customer Success](/video/creating-scalable-ai-factories-built-for-customer-success)

**Published:** January 21, 2026

Accelerate your Time to First Token (TTFT). Build a scalable AI factory to optimize your AI infrastructure investment and guarantee enterprise ROI.

## [How Memory-First Architecture Solves AI Inference Challenges](/video/how-memory-first-architecture-solves-ai-inference-challenges)

**Published:** January 20, 2026

Shatter the memory wall. See how optimizing GPU memory bandwidth reduces inter-token latency (ITL) to eliminate bottlenecks in production AI inference.

## [Architecting AI Infrastructure for the Age of Reasoning](/video/architecting-ai-infrastructure-for-the-age-of-reasoning)

**Published:** January 20, 2026

Maximize your AI infrastructure investment. Solve storage latency to scale your AI factory and eliminate bottleneck failures during production inference.

## [Driving Faster Time to Production for AI Inference](/video/driving-faster-time-to-production-for-ai-inference)

**Published:** January 20, 2026

Resolve GPU storage bottlenecks. Optimize your AI infrastructure investment and deploy reference architectures to speed real-world AI factory inference.

## [Why Infrastructure Will Catalyze AI with Meta, Lambda, & Silicon Data](/video/why-infrastructure-will-catalyze-ai-with-meta-lambda-silicon-data)

**Published:** January 15, 2026

Maximize your AI infrastructure investment. Solve GPU capacity limits to eliminate bottlenecked models and scale your enterprise AI factory fast.

## [Inside the AI Capacity Crunch: Solving Latency, Memory Limits, and Multi-Agent Scaling | VentureBeat AI Impact Tour NYC](/video/inside-the-ai-capacity-crunch-solving-latency-memory-limits-and-multi-agent-scaling-or)

**Published:** November 20, 2025

Solve the GPU memory bottleneck. Discover how optimizing pre-fill latency and KV cache scaling prevents multi-agent application failures.

## [Efficient Infrastructure Design is Transforming the Future of AI](/video/efficient-infrastructure-design-is-transforming-the-future-of-ai)

**Published:** November 5, 2025

Stop wasting expensive GPU utilization. Streamline your AI infrastructure investment to optimize tokenomics and eliminate model latency today.

## [The Inference Era: Building Scalable Data Infrastructure for AI with Nand Research](/video/the-inference-era-building-scalable-data-infrastructure-for-ai-with-nand-research)

**Published:** November 4, 2025

Stop GPU idle time. Optimize scale-out storage costs and TCO for AI storage to eliminate bottlenecks and dramatically maximize enterprise Inference ROI.

## [AI Inference, Agent Swarms, and Token Economics | Val Bercovici at VentureBeat AI Impact Tour](/video/ai-inference-agent-swarms-and-token-economics-val-bercovici-at-venturebeat-ai-impact-tour)

**Published:** October 14, 2025

Val Bercovici, Chief AI Strategy Officer at WEKA, explains the AI inference capacity crisis, rising token costs, reasoning models, and why agent swarms reshape infrastructure.

## [It's Time to Put Your Data in the Fast Lane](/video/its-time-to-put-your-data-in-the-fast-lane)

**Published:** October 7, 2025

Get rid of latency bottlenecks and metadata overload. Power your data to move faster and smarter with NeuralMesh.

## [Inside NeuralMesh by WEKA: A Containerized Microservices Architecture](/video/inside-neuralmesh-by-weka-a-containerized-microservices-architecture)

**Published:** June 27, 2025

Introducing NeuralMesh™️ by WEKA®️. The world’s only storage system purpose-built for AI. What it is, why it matters, and how it helps AI leaders overcome the biggest barriers to scale and speed.

## [Introducing NeuralMesh™ by WEKA®](/video/introducing-neuralmesh-by-weka)

**Published:** June 18, 2025

Get an inside look at NeuralMesh by WEKA —the world’s only storage system purpose-built for AI.

## [AI Infra Summit: WEKA CEO Liran Zvibel on Tokenomics, Scaling, and the Future of Infrastructure](/video/ai-infra-summit-weka-ceo-liran-zvibel-on-tokenomics-scaling-and-the-future-of-infrastructure)

**Published:** May 30, 2025

WEKA CEO Liran Zvibel explains how tokenomics and memory-efficient infrastructure can future-proof enterprise AI at scale.

## [AI Infra Summit Recap: Tokenomics. Latency. Survival Economics for GenAI](/video/ai-infra-summit-val-bercovici-keynote-tokenomics-latency-survival-economics-for-genai)

**Published:** May 30, 2025

Explore WEKA's Val Bercovici’s take on AI tokenomics, memory limits, and economic survival in the GenAI era at the AI Infra Summit.

## [AI Infra Summit: Shimon Ben-David, WEKA and Nave Algarici, NVIDIA discuss what is next in AI Infrastructure](/video/ai-infra-summit-shimon-ben-david-weka-and-nave-algarici-nvidia-discuss-what-is-next-in-ai-infrastructure)

**Published:** May 30, 2025

WEKA and NVIDIA share how to scale AI inferencing with real-time data and GPU efficiency at the AI Infra Summit.

## [How WEKA’s Augmented Memory Grid™ brings petabytes of persistent storage for KV Cache](/video/how-wekas-augmented-memory-grid-brings-petabytes-of-persistent-storage-for-kv-cache)

**Published:** May 15, 2025

Learn how Augmented Memory Grid radically improves the economics and performance of AI inference.

## [Blueprint for Supercharging LLM Inference With "PagedAttention over RDMA"](/video/blueprint-for-supercharging-llm-inference-with-pagedattention-over-rdma)

**Published:** April 18, 2025

PagedAttention over RDMA" (PAoR) revolutionizes large language model serving by addressing key-value (KV) cache challenges with RDMA networking

## [AI Explained: Understanding Multicloud GenAI Workloads](/video/ai-demystified-episode-06-ai-explained-understanding-multicloud-genai-workloads)

**Published:** April 1, 2025

Discover the primary obstacles in multicloud AI workloads, from high transfer costs to data duplication issues, and learn how to mitigate issues.

## [AI Explained: Monitoring GenAI Workloads with NVIDA’s NeMo](/video/ai-demystified-episode-05-ai-explained-monitoring-genai-workloads-with-nvidas-nemo)

**Published:** April 1, 2025

Dive into best practices to keep your AI training running smoothly when using NVIDIA's NeMo.

## [AI Explained: Vector Databases and AI Performance in RAG Pipelines](/video/ai-demystified-episode-04-ai-explained-vector-databases-and-ai-performance-in-rag-pipelines)

**Published:** April 1, 2025

Learn about the role of vector databases in storing and querying embeddings for RAG systems, and more.

## [AI Token Efficiency](/video/ai-token-efficiency)

**Published:** April 1, 2025

WEKA's Val Bercovici provides an overview of AI tokenomics and the implications for AI-first organizations.

## [WEKA Augmented Memory](/video/weka-augmented-memory)

**Published:** March 18, 2025

Augmented Memory Grid in NeuralMesh capability extends large, persistent memory capabilities to optimize inference pipelines.

## [NeuralMesh vs. Amazon FSX for Lustre](/video/weka-data-platform-vs-amazon-fsx-for-lustre)

**Published:** February 27, 2025

A comparative analysis between NeuralMesh and Amazon FSX for Lustre, focusing on metadata performance and directory navigation using the “ls” command.

## [AI Explained: Checkpointing in LLMs and the Trade-Offs between Reliability and Performance](/video/ai-demystified-episode-03-ai-explained-checkpointing-in-llms-and-the-trade-offs-between-reliability-and-performance)

**Published:** February 26, 2025

Checkpointing is a crucial process that ensures AI models, including large language models (LLMs), continue training efficiently and without setbacks.

## [AI Explained: How Retrieval-Augmented Generation (RAG) Transforms Large Language Models (LLMs)](/video/ai-demystified-episode-02-how-rag-transforms-llms)

**Published:** February 25, 2025

Retrieval-Augmented Generation (RAG) is a groundbreaking technique in natural language processing (NLP) that combines the power of information retrieval with generative AI (GenAI) models.

## [Building GenAI Infrastructure: 5 Key Features of NVIDIA NIM](/video/ai-demystified-episode-01-building-genai-infrastructure-5-key-features-of-nvidia-nim)

**Published:** February 25, 2025

NVIDIA’s NeMo Inference Microservices (NIM) platform is a solution for serving a wide range of AI models, in the areas of natural language processing (NLP), computer vision, and speech recognition.

## [Picture Shop Partners with WEKA](/video/picture-shop-partners-with-weka)

**Published:** December 20, 2024

Picture Shop is a global post-production company that works on movies, episodic, commercials, and unscripted entertainment. NeuralMesh enabled Picture Shop’s finishing teams to work efficiently in parallel.

## [Dead & Company Dead Forever at Sphere](/video/dead-company-dead-forever-at-sphere)

**Published:** October 17, 2024

NeuralMesh drives creativity and optimizes technology with a performant, flexible solution for Dead & Company’s Dead Forever residency at Sphere.

## [The Evolution of AI <span Trends and Emerging Challenges</span>](/video/the-evolution-of-ai)

**Published:** September 11, 2024

Watch this video with 451 Research Senior Research Analyst, Alex Johnston, sharing research and analysis across key AI trends in 2024.

## [Six Five on the Road Featuring CTO Shimon Ben David](/video/six-five-on-the-road-featuring-cto-shimon-ben-david)

**Published:** July 29, 2024

WEKA Chief Technology Officer Shimon Ben David joins Six Five On the Road with hosts Dave Nicholson and Alastair Cooke for a conversation on how WEKA’s solutions are designed to keep storage technology in pace with the rapidly evolving demands of today’s AI workloads.

## [The Impossibles: Vera C. Rubin Observatory](/video/the-impossibles-vera-c-rubin-observatory)

**Published:** June 28, 2024

See how the Vera C. Rubin Observatory, located on Cerro Pachón in Chile, is transforming our understanding of the universe through a 10-year survey of the southern hemisphere known as the Legacy Survey of Space and Time (LSST) initiative.

## [NYSE Floor Talk: WEKA Officially Becomes A Unicorn!](/video/nyse-floor-talk-weka-officially-becomes-a-unicorn)

**Published:** May 23, 2024

WEKA President Jonathan Martin joins Judy Shaw on NYSE Floor Talk to discuss WEKA's latest funding raise, unicorn status, and the future of AI infrastructure.

## [The Impossibles: U2 at Sphere](/video/the-impossibles-u2-at-sphere)

**Published:** April 23, 2024

## [WEKApod: The World’s Fastest AI Data Infrastructure](/video/wekapod-the-worlds-fastest-ai-data-infrastructure)

**Published:** April 15, 2024

## [Introducing WEKApod™](/video/introducing-wekapod)

**Published:** March 18, 2024

## [See How NexGen Cloud Is Democratizing AI With WEKA](/video/nexgen-cloud-is-democratizing-ai-with-gpu-cloud-services-powered-by-weka)

**Published:** January 31, 2024

## [WEKA Helps Applied Digital Serve Continuous Data to the Fastest GPUs on the Planet](/video/applied-digital-uses-weka-to-serve-continuous-data-up-to-the-fastest-gpus-on-the-planet)

**Published:** November 16, 2023

## [Don’t Let Lots of Small Files Slow Down Your AI Projects](/video/dont-let-lots-of-small-files-slow-down-your-ai-projects)

**Published:** November 6, 2023

One of the challenges of training AI models is using data sets that contain lots of small files (LOSF). Despite the unassuming acronym, these can be quite a challenge for legacy storage. NeuralMesh easily handles all the files needed to train AI models no matter the number or size.

## [WEKA Data Reduction](/video/weka-data-reduction)

**Published:** November 1, 2023

A demonstration of WEKA data reduction technology including how to configure it and space savings from using it.

## [Sloths Sleep Over 50% of the Day ... What About Your GPU?](/video/sloths-sleep-over-50-of-the-day-what-about-your-gpu)

**Published:** April 16, 2023

## [TRE ALTAMIRA Chose NeuralMesh for Geospatial Research](/video/tre-altamira-chose-distributed-file-system-innovator-wekaio-for-geospatial-research)

**Published:** April 15, 2023

Tre Altamira chose NeuralMesh for their geospatial research, find out why in this informative video showcasing the benefits of the distributed storage system.

## [Clovertex and WEKA: Faster Drug Discovery Design with NeuralMesh on AWS](/video/clovertex-and-weka-faster-drug-discovery-design-with-weka-on-aws)

**Published:** April 14, 2023

## [Samsung Fireside Chat](/video/samsung-fireside-chat)

**Published:** April 13, 2023

Join our fireside chat with Samsung and learn how NeuralMesh is revolutionizing data storage for the digital age.

## [23andMe on AWS with NeuralMesh by WEKA](/video/23andme-on-aws-with-weka)

**Published:** April 13, 2023

Watch to see how 23andMe leveraged the power of NeuralMesh and AWS to accelerate their genomic data analysis workflows.

## [Delivering Unheard of Performance on OCI with WEKA](/video/delivering-unheard-of-performance-on-oci-with-weka)

**Published:** April 13, 2023

## [Preymaker Runs A Cloud-Native VFX Studio Managed by One Person on NeuralMesh](/video/preymaker-runs-a-cloud-native-vfx-studio-managed-by-one-person-on-weka)

**Published:** April 13, 2023

Preymaker is using the NeuralMesh on AWS to share data across different applications and OS while retaining visibility and control of cloud resources with predictable costs.

## [Data Portability Between Clouds Demo](/video/data-portability-between-clouds-demo)

**Published:** April 12, 2023

Watch our video about data portability between clouds with NeuralMesh. Discover the potential of our innovative technology today.

## [Run the Impossible. Anywhere.](/video/run-the-impossible-anywhere)

**Published:** April 12, 2023

## [Are Your GPUs on a Catnap? Discover Accelerated Purrrfection with NeuralMesh](/video/are-your-gpus-on-a-catnap-discover-accelerated-purrrfection-with-weka)

**Published:** April 12, 2023

If your GPU accelerated workloads are experiencing data stalls and management challenges, this could lead to your GPUs being underutilized. NeuralMesh accelerates every step of your AI data pipeline, leading to lower epoch time, and faster times to insight.

## [Festival of Genomics: Fireside Chat with Genomics England](/video/festival-of-genomics-fireside-chat-with-genomics-england)

**Published:** April 12, 2023

Watch the Festival of Genomics Fireside Chat with Genomics England and learn  how they use NeuralMesh to meet future capacity scaling requirements and to achieve the highest performance to its DNA pipeline.

## [Untold Studios Chooses NeuralMesh for AI on AWS for Hybrid Cloud Storage](/video/untold-studios-chooses-weka-the-data-platform-for-ai-on-aws-for-hybrid-cloud-storage)

**Published:** April 11, 2023

Untold Studios uses NeuralMesh because it is cloud-native, simple, eliminates management overhead, and scales with their business needs.

## [Carbon VFX Breaks Down Barriers to Creativity with NeuralMesh and AWS](/video/carbon-vfx-breaks-down-barriers-to-creativity-with-weka-and-aws)

**Published:** March 7, 2023

Watch how Carbon VFX leveraged the power of NeuralMesh and AWS to break down barriers to creativity and accelerate their workflows.
