Your Storage is the Bottleneck. Not Your GPUs.
AI teams on AWS are burning GPU dollars waiting on FSx for Lustre. There’s a better architecture — and it doesn’t require leaving AWS.
Why FSx Can’t Keep Up With What You’re Building
Cast AI clocked it in April: average GPU utilization across customer clusters is 5%. The gap isn’t the scheduler or the model — it’s storage. FSx for Lustre was built for HPC scratch space, not sustained training, concurrent inference, or multi-tenant GPU clusters with per-team SLAs.
Where FSx hits the wall:
- Throughput-capacity coupling forces overprovisioning you don’t need
- Single-AZ default leaves multi-day training jobs exposed to zone failures
- No persistent KV cache means TTFT degrades under concurrent inference load
- No native QoS means one team’s training starves another’s serving
Stability AI ran the experiment. Same AWS footprint. Different storage layer.
FSx for Lustre was the ceiling. NeuralMeshâ„¢ removed it.
95% cost reduction
Storage spend dropped from $5M to $1M annually — without touching a single GPU or moving regions.
93% GPU utilization
A 400-node cluster hitting 93% utilization. Lustre was leaving most of that GPU spend on the floor.
35% faster training
Three weeks cut from a 60-day training cycle. Not from better models — from storage that stopped getting in the way.
15x more capacity, same cost
Same AWS instances. Same NVMe. Fifteen times more usable capacity at 80% of the previous cost.
Stability AI ran the experiment. Same AWS footprint. Different storage layer.
FSx for Lustre was the ceiling. NeuralMeshâ„¢ removed it.
95% Cost Reduction
Storage spend dropped from $5M to $1M annually — without touching a single GPU or moving regions.
93% GPU Utilization
A 400-node cluster hitting 93% utilization. Lustre was leaving most of that GPU spend on the floor.
35% Faster Training
Three weeks cut from a 60-day training cycle. Not from better models — from storage that stopped getting in the way.
15x More Capacity, Same Cost
Same AWS instances. Same NVMe. Fifteen times more usable capacity at 80% of the previous cost.
Stop guessing. See your numbers.
Want to benchmark your current setup? Talk to a WEKA expert about your workload, your AWS bill, and where the bottlenecks are.
Performance results based on Stability AI production deployment on AWS, 2025. Cast AI GPU utilization data from April 2026 industry study. Results vary by workload and configuration. WEKA NeuralMesh runs natively on AWS — no migration off cloud required.