Webinar
Does Agentic AI Break the Case for GPU-Only Inference?
Join us as SemiAnalysis moderates a live fireside discussion between WEKA and Oracle on where CPU inference actually belongs in production AI, with a live demo of real-world agent swarm performance.
Why attend:
- Watch a live demo of an inference workload on OCI + WEKA and see the KV cache bottleneck and the fix in action, not described in a slide
- Hear SemiAnalysis push WEKA and Oracle on where the CPU/GPU line actually sits — this isn’t your typical scripted vendor session
- Get straight answers on where your inference budget is actually going
What you’ll leave with:
A framework for evaluating your own inference stack — where CPU should be doing the work, where GPU still wins, and how to tell the difference in your own environment. Not a takeaway deck; a way of thinking about the next architecture decision in front of you.
Agent sessions aren’t one continuous GPU pass, they’re dozens of model steps interrupted by CPU-side tool calls, retrieval, and code execution. Every one of those gaps risks evicting the GPU’s KV cache, forcing an expensive recompute. Recent industry benchmarking, including NVIDIA’s own analysis of agentic workloads, points to this exact failure point. This session covers what it means for your architecture on OCI, live.
Agenda:
- Welcome & Framing
- Fireside Discussion, moderated by SemiAnalysis
- Live Demonstration
- Open Discussion and Q&A
Meet Our Speakers
Val Bercovici
Chief AI Officer, WEKA
Manu Mishra
Principal Solutions Architect, Oracle
Max Kan
Tokenomics Technical Lead, SemiAnalysis
Watch On-Demand
Please register using your business email address. Registrations from personal email domains (e.g., Gmail, Yahoo, etc.) will not be confirmed.