This episode explores SnapStream, a technique from SambaNova Systems for compressing KV caches during long-sequence LLM decoding on dataflow accelerators, demonstrated at production scale with a 671-billion-parameter DeepSeek-R1 deployment running 128K-token context at over 1,800 tokens per second. The discussion covers why established training-free KV cache eviction methods like SnapKV and StreamingLLM have struggled to reach real deployments despite promising accuracy results: continuous batching makes it unclear when to trigger compression across requests at different lifecycle stages, and static-graph compilers used by dataflow accelerators can't easily accommodate the dynamic, variable-shaped operations that standard compression implementations rely on. It explains how SnapStream fuses SnapKV's attention-based token selection with StreamingLLM's sink-plus-sliding-window approach into a single fixed-size cache, splitting sequences into sink tokens, recent tokens, and a compressed middle section during prefill. The conversation is grounded in fundamentals—clarifying the prefill/decode split, why decode is memory-bound, and what makes dataflow accelerators architecturally different from GPUs—making it accessible to listeners unfamiliar with KV cache mechanics while still delivering a specific, hardware-grounded engineering story rather than a purely algorithmic one. Sources: 1. SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators — Jonathan Li, Nasim Farahini, Evgenii Iuliugin, Magnus Vesterlund, Christian Häggström, Guangtao Wang, Shubhangi Upasani, Ayush Sachdeva, Rui Li, Faline Fu, Chen Wu, Ayesha Siddiqua, John Long, Tuowen Zhao, Matheen Musaddiq, Håkan Zeffer, Yun Du, Mingran Wang, Qinghua Li, Bo Li, Urmish Thakker, Raghu Prabhakar, 2025 http://arxiv.org/abs/2511.03092 2. Plasticine: A Reconfigurable Architecture For Parallel Patterns — Raghu Prabhakar, Yaqi Zhang, David Koeplinger, Matt Feldman, Tian Zhao, Stefan Hadjis, Ardavan Pedram, Christos Kozyrakis, Kunle Olukotun, 2017 https://scholar.google.com/scholar?q=Plasticine%3A+A+Reconfigurable+Architecture+For+Parallel+Patterns 3. SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts — Raghu Prabhakar and SambaNova Systems architecture team, 2024 https://scholar.google.com/scholar?q=SN40L%3A+Scaling+the+AI+Memory+Wall+with+Dataflow+and+Composition+of+Experts 4. Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads — Dennis Abts, Jonathan Ross, Jonathan Sparling, Mark Wong-VanHaren, et al. (Groq), 2020 https://scholar.google.com/scholar?q=Think+Fast%3A+A+Tensor+Streaming+Processor+%28TSP%29+for+Accelerating+Deep+Learning+Workloads 5. In-Datacenter Performance Analysis of a Tensor Processing Unit — Norman P. Jouppi et al. (Google), 2017 https://scholar.google.com/scholar?q=In-Datacenter+Performance+Analysis+of+a+Tensor+Processing+Unit 6. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized LLM Serving — Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, H. Zhang, 2024 https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+LLM+Serving 7. Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference — J. Tang, Y. Zhao, K. Zhu, G. Xiao, B. Kasikci, S. Han, 2024 https://scholar.google.com/scholar?q=Quest%3A+Query-Aware+Sparsity+for+Efficient+Long-Context+LLM+Inference 8. InfLLM: Training-Free Long-Context Extrapolation with an Efficient Context Memory — C. Xiao, P. Zhang, X. Han, G. Xiao, Y. Lin, Z. Zhang, Z. Liu, M. Sun, 2024 https://scholar.google.com/scholar?q=InfLLM%3A+Training-Free+Long-Context+Extrapolation+with+an+Efficient+Context+Memory 9. DeepSeek-V3.2-Exp: Boosting Long-Context Efficiency with DeepSeek Sparse Attention — DeepSeek-AI, 2025 https://scholar.google.com/scholar?q=DeepSeek-V3.2-Exp%3A+Boosting+Long-Context+Efficiency+with+DeepSeek+Sparse+Attention 10. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study — S. Eyuboglu, R. Ehrlich, S. Arora, N. Guha, D. Zinsley, E. Liu, W. Tennien, A. Rudra, J. Zou, A. Mirhoseini, C. Re, 2025 https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+General-Purpose+Long+Context+Representations+via+Self-Study 11. LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction (SAGE-KV) — G. Wang, S. Upasani, C. Wu, D. Gandhi, J. Li, C. Hu, B. Li, U. Thakker, 2025 https://scholar.google.com/scholar?q=LLMs+Know+What+to+Drop%3A+Self-Attention+Guided+KV+Cache+Eviction+%28SAGE-KV%29 Interactive Visualization: SnapStream: Taming KV-Cache Eviction for Dataflow Accelerators