Webinar

Does Agentic AI Break the Case for GPU-Only Inference?

Join us as SemiAnalysis moderates a live fireside discussion between WEKA and Oracle on where CPU inference actually belongs in production AI, with a live demo of real-world agent swarm performance.

Why attend:

  • Watch a live demo of an inference workload on OCI + WEKA and see the KV cache bottleneck and the fix in action, not described in a slide
  • Hear SemiAnalysis push WEKA and Oracle on where the CPU/GPU line actually sits — this isn’t your typical scripted vendor session
  • Get straight answers on where your inference budget is actually going

What you’ll leave with:

A framework for evaluating your own inference stack — where CPU should be doing the work, where GPU still wins, and how to tell the difference in your own environment. Not a takeaway deck; a way of thinking about the next architecture decision in front of you.

Agent sessions aren’t one continuous GPU pass, they’re dozens of model steps interrupted by CPU-side tool calls, retrieval, and code execution. Every one of those gaps risks evicting the GPU’s KV cache, forcing an expensive recompute. Recent industry benchmarking, including NVIDIA’s own analysis of agentic workloads, points to this exact failure point. This session covers what it means for your architecture on OCI, live.

Agenda:

  • Welcome & Framing
  • Fireside Discussion, moderated by SemiAnalysis
  • Live Demonstration
  • Open Discussion and Q&A

Meet Our Speakers

Speakers

Val Bercovici

Chief AI Officer, WEKA

Speakers

Manu Mishra

Principal Solutions Architect, Oracle

Speakers

Max Kan

Tokenomics Technical Lead, SemiAnalysis

Watch On-Demand

Please register using your business email address. Registrations from personal email domains (e.g., Gmail, Yahoo, etc.) will not be confirmed.