# Blueprint for Supercharging LLM Inference With "PagedAttention over RDMA"

**Published:** April 18, 2025

![Blueprint for Supercharging LLM Inference With "PagedAttention over RDMA"](https://cdn.sanity.io/images/ult5g8gw/production/355f68b0903c8cff111cfcf84d397a5e6a861523-960x540.jpg)

**Watch:** https://fast.wistia.net/embed/iframe/wrjhzilfq9

PagedAttention over RDMA" (PAoR) revolutionizes large language model serving by addressing key-value (KV) cache challenges with RDMA networking

## Commentary

Learn how "PagedAttention over RDMA" (PAoR) revolutionizes large language model serving by addressing key-value (KV) cache challenges with RDMA networking and NeuralMesh. This session showcases seamless integration with vLLM and TensorRT-LLM, enabling faster inference with reduced latency and increased throughput across multi-node environments.

## Related Videos

- [Inference Is Eating Memory, and Tokens Now Run AI Economics](/video/inference-is-eating-memory-and-tokens-now-run-ai-economics)
- [Build for What's Next in AI with NeuralMesh](/video/build-for-what-s-next-in-ai-with-neuralmesh)
- [WEKApod: Maximum Capacity and Performance Density](/video/wekapod-maximum-capacity-and-performance-density)
