Keeping GPUs Busy Is the Real AI Infrastructure KPI
Idle GPUs stall AI workloads. Hear HTX's Chief Innovation Officer explain how optimized infrastructure shortens the data circuit to maximize GPU utilization.
What makes GPU infrastructure investment pay off?
Betsy Chernoff: I’m Betsy Chernoff with WEKA.
NG Pan Yong: I’m NG Pan Yong, Chief Innovation Officer for HTX, and thanks WEKA for being here.
Betsy: Thank you for letting me talk to you for a little bit, Pan Yong. When you guys are looking to invest in infrastructure, what are some of the performance metrics? What are some of the things that get you most excited for it?
Pan Yong: So maybe the way to frame it is, the investment that we make for GPUs is quite significant. What we want to do is maximize the investment we place in the GPU, and every time the GPU is idle, not doing something, it’s a real concern for us. So in terms of a technology partner, for example with WEKA, we are looking for a solution that gives us low latency and high throughput, be it for checkpointing during training, as part of KV caching for inference, or being the core platform for us to host our AI models for inference.
Betsy: Sure. So it’s fair to say that keeping your GPUs as hydrated as possible is probably one of the most important things?
Pan Yong: My most important KPI. So there are two KPIs: One is to buy GPUs, and the second is to keep them really busy.
Betsy: That makes sense.
Why storage became mission-critical for AI performance
Betsy: So when we talk about this world of AI, one of the things everybody is talking about — even us right now — we’re talking about GPUs and how important they are. But the world has really changed from just focusing on compute. Now we talk about other layers as well, like storage and networking. Oftentimes we’ve heard that storage was maybe a second-class citizen, if you will, compared to compute. But in this day and age it’s obviously a little bit different. I wonder how you guys see it, and what you’ve seen change in the world of data as well.
Pan Yong: I think you’re absolutely right. If I take one step back, traditionally, when it was CPU-centric, it didn’t matter as much. We know the business stays up the stream. We can deal with slower-performance storage, like maybe a SAN and things like that. The moment we get into GPUs, it’s really accelerated computing with lots of parallel processing. Having that integrated stack, optimizing the networking, optimizing the connectivity between storage and GPU becomes important.
The way you can look at it is this: You have lots of data, and the GPU is the brain. Can we shorten the circuit between the data and the brain to optimize our inference and performance? It’s super important for us.
Betsy: And it’s definitely the moment.
What to look for in an AI storage partner
Betsy: When you’re looking overall for a storage architecture, what are some of the pieces that are most important for you and your team to look for?
Pan Yong: I think there is a technology element: performance at the flash level, performance at the software stack level, performance at the software level.
But to be honest, the most important thing for us is the innovation of the partners we work with. Are you looking beyond just being the storage company? Are you saying, “I’m going to be at the frontier of AI”? I’m understanding: How are we going to do distributed inferencing? How are we going to optimize performance?” Does the organization we work with understand the ebbs and flows of where the innovation is going?
And we’re happy WEKA as a partner is at the frontier of doing some of these things, and that’s what we’re looking for in a partner on this journey for AI.
Betsy: Wonderful, and we’re thrilled to be a partner with you as well. Thank you so much for your time.
Pan Yong: Thank you.
Featured Speakers
- Betsy ChernoffPrincipal Product Marketing ManagerWEKA
- NG Pan YongChief Innovation OfficerHTX
What's Next
Scale Production AI Faster with NeuralMesh
Your models aren't slow. Your data is. Fix AI bottlenecks with high-throughput infrastructure.


