# WEKA® NeuralMesh Axon™

**Make Your AI Infrastructure Work Harder**

Turn underused resources inside your GPU fleet into high-performance shared storage closer to compute.

## There’s More Inside Your GPU Fleet

### Converged Compute and Storage

**Bring Shared Storage Inside Your GPU Fleet**

Run WEKA® NeuralMesh™ directly across GPU servers to put high-performance shared storage inside the compute infrastructure. Allocated CPU, memory, network, and NVMe resources power storage services, bringing data closer to GPU compute while putting more server resources to productive use.

- [Read the Architecture](/resources/white-paper/wekaio-architectural-whitepaper)

### WEKA Augmented Memory Grid

**Unified Storage and Memory**

Pair NeuralMesh Axon with WEKA Augmented Memory Grid™ to unify storage and memory for inference. NeuralMesh Axon puts NeuralMesh inside the GPU infrastructure where Augmented Memory Grid can extend GPU memory into high-speed, persistent storage for key-value (KV) cache. Unify storage and memory to preserve reusable context, reduce recomputation, and keep GPUs generating tokens.

- [See Augmented Memory Grid](/product/augmented-memory-grid/)

### Local NVMe Aggregation

**Turn Stranded Flash Into Shared Storage**

NeuralMesh Axon aggregates local NVMe across GPU servers into shared NeuralMesh storage. Flash isolated on individual servers becomes productive capacity available across the cluster, helping you get more value from storage resources you already own.

- [Read the Architecture](/resources/white-paper/wekaio-architectural-whitepaper)

### Unified Namespace

**Make Local Data Available Across the Fleet**

NeuralMesh Axon creates a unified namespace across the NeuralMesh Axon environment, turning server-local storage into shared infrastructure. AI workloads can access data across the cluster regardless of which server holds the underlying flash.

- [Unified Namespace](/resources/white-paper/wekaio-architectural-whitepaper)

### Resilient Failure Domains

**Ensure Data Resilience Across the Fleet**

NeuralMesh Axon distributes data across participating GPU servers instead of depending on an individual server or drive. NeuralMesh failure domains enable local resources to operate as resilient shared storage when individual infrastructure components become unavailable.

### Resource Management

**Control Exactly What Storage Uses**

NeuralMesh Axon allocates specific CPU, memory, network, and NVMe resources to NeuralMesh storage services alongside AI workloads. Keep storage performance predictable while maintaining control over the resources available to AI workloads.

- [Unlock Your GPUs](/resources/solution-brief/weka-neuralmesh-axon-solution-brief)

## Key Stats

- **93%** *GPU Utilization* — Keep GPUs productive with an architecture designed to make better use of the infrastructure inside the GPU fleet.
- **10x** *more concurrent users* — Pair NeuralMesh Axon with WEKA Augmented Memory Grid and serve up to 10x more concurrent users for inference.
- **30 GBps** *Throughput per GPU Server* — Deliver high-performance shared storage from within the GPU fleet.

## Customer Story

> Embedding WEKA's NeuralMesh Axon into our GPU servers enabled us to maximize utilization and accelerate every step of our AI pipelines. The performance gains have been game-changing: Inference deployments that used to take five minutes can occur in 15 seconds, with 10 times faster checkpointing. Our team can now iterate on and bring revolutionary new AI models, like North, to market with unprecedented speed.

— Autumn Moulder, Cohere

## Frequently Asked Questions

### What is NeuralMesh Axon?

NeuralMesh Axon is a converged deployment of NeuralMesh that runs directly across GPU servers. It uses allocated CPU, memory, network, and local NVMe resources to deliver shared, high-performance storage, bringing storage closer to GPU compute while making more efficient use of infrastructure already inside the GPU fleet.

### How do NeuralMesh Axon and Augmented Memory Grid work together?

NeuralMesh Axon puts NeuralMesh storage inside GPU infrastructure. Augmented Memory Grid uses NeuralMesh as a persistent, NVMe-backed token warehouse for KV cache, extending the memory hierarchy for inference.

### When should I use converged versus dedicated AI storage?

Use converged storage when infrastructure efficiency and proximity to GPU compute are priorities. Use dedicated storage when resource isolation and independent scaling are priorities. Compare NeuralMesh deployment models.

### How can storage share resources with GPU workloads?

NeuralMesh Axon allocates specific CPU, memory, network, and NVMe resources to storage services, isolating them from resources assigned to AI workloads. Explore the NeuralMesh Axon architecture.

### What AI workloads benefit from converged storage?

Converged storage benefits data-intensive training and inference workloads. NeuralMesh Axon brings shared storage close to GPU compute for training, checkpointing, model loading, and inference at scale. See NeuralMesh Axon for training and inference.

### How can converged storage improve AI infrastructure efficiency?

Converged storage puts underused GPU server resources to productive use. NeuralMesh Axon uses local NVMe and allocated server resources as shared storage, reducing the need for separate storage infrastructure. See how NeuralMesh Axon works.

## Related Resources

- [WEKA® NeuralMesh™ Architecture White Paper](/resource/wekaio-architectural-whitepaper) (White Paper)
- [NeuralMesh Axon: Unlock The Full Potential of Your GPUs](/resource/weka-neuralmesh-axon-solution-brief) (Solution Brief)
- [Center for AI Safety](/customers/center-for-ai-safety-cais/) (Case Study)
