1 Million Tokens Per Second on Kubernetes, with Federico Iezzi

GPU inference throughput depends on more than accelerator generation or count.

Memory bandwidth, model parallelism, cache configuration, and the load generator itself all influence measured throughput.

Federico Iezzi, Customer Engineer at Google Cloud, explains how his team achieved 1 million output tokens per second using Qwen 3.5 27B, vLLM, GKE Autopilot, and NVIDIA B200 GPUs.

The discussion covers:

  1. Why memory bandwidth limits decode performance

  2. How Federico chose between tensor and data parallelism

  3. What changed after enabling multi-token prediction and reducing the KV cache footprint with FP8 quantization.

Sponsor

This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.

More info

  • Find all the links and info for this episode here: https://ku.bz/1xD9Md0mb

  • Interested in sponsoring an episode? Learn more.