Kubernetes Bytes

Ryan Wallner & Bhavin Shah

Kubernetes Bytes is a podcast bringing you the latest from the world of cloud native data management. Hosts Ryan Wallner and Bhavin Shah come to you from Boston, Massachusetts with experienced backgrounds in cloud-native tech. They'll be sharing their thoughts on recent cloud native news and talking to industry experts about their experiences and challenges managing the wealth of data in today's cloud-native ecosystem.

  1. 3h ago

    What if Kubernetes Could Resume Instead of Restart? | Checkpoint / Restore Explained

    Kubernetes usually handles failures and rescheduling by restarting workloads. But what if Kubernetes could save the exact running state of a workload, move it somewhere else, and resume from where it stopped?   In this episode of Kubernetes Bytes, Bhavin sits down with Radostin Stoyanov, maintainer of CRIU (Checkpoint/Restore In Userspace) and a contributor helping lead the Kubernetes Checkpoint/Restore Working Group.   They explore how checkpoint/restore works, why it is becoming increasingly relevant for Kubernetes and AI workloads, and how the community is working toward pod-level checkpoint and restore.   The conversation covers how checkpointing can help reduce expensive AI inference cold starts, improve fault tolerance for large GPU training jobs, enable faster model swapping, and potentially influence future Kubernetes scheduling and preemption workflows.  They also discuss the challenges of checkpointing distributed applications, GPU memory, compression, security, encryption, and integrating checkpoint/restore directly into Kubernetes APIs.   Topics include:  What checkpoint/restore actually saves Application-level vs. infrastructure-level checkpointing CRIU and Kubernetes Why AI inference workloads can take minutes to initialize Saving CPU and GPU execution state  Pod-level vs. container-level checkpointing  Kubernetes Checkpoint/Restore Working Group  Distributed checkpoint/restore  Fault tolerance for large GPU training jobs  Accelerating model startup and model swapping  Memory compression for large checkpoints Kubernetes APIs and controllers for checkpoint/restore  Security risks of checkpoint files  Encryption and validation  Scheduler preemption and workload migration  Checkpoint/restore for AI-agent sandboxes   If you work with Kubernetes, GPUs, AI infrastructure, inference, distributed training, or scheduling, this episode provides a look at a Kubernetes capability that could become increasingly important as workloads become more expensive to restart.   Subscribe to Kubernetes Bytes for conversations about Kubernetes, cloud native infrastructure, AI, platform engineering, security, and the CNCF ecosystem.

  2. Sep 15

    Your AI Agent Has Too Much Access | Fixing Agent Identity

    AI agents can do far more than answer questions — they can call tools, access data, execute actions, and act on behalf of users. But once agents start taking actions, an important question emerges: who is the agent actually acting as, and what should it be allowed to do? In this episode of Kubernetes Bytes, Bhavin sits down with Maia Iyer, Research Software Engineer at IBM Research, to explore identity, authentication, authorization, and zero trust for AI agents. They discuss why static API keys and simply passing user credentials to agents can create security problems, and how existing standards such as OAuth 2.0, SPIFFE, SPIRE, and Keycloak can help establish verifiable user and workload identities. Maia also explains the work her team is doing around Rosso CTL and Cortex, including using a runtime layer around agents to handle identity, observe inbound and outbound interactions, apply guardrails, and explore intent-based access control. Topics covered include: • The building blocks of agentic applications• Security anti-patterns for AI agents• Static credentials and over-permissioned agents• User identity vs. workload identity• OAuth 2.0 and delegated identity• SPIFFE and SPIRE• Keycloak• Rosso CTL and Cortex• Intent-based access control• Prompt injection and unexpected tool calls• MCP and agent-to-tool communication• Stateful AI agents• Safely running agents that generate and execute code If you're building AI agents on Kubernetes or thinking about how identity and authorization need to evolve for agentic systems, this episode provides a practical mental model for getting started. Subscribe to Kubernetes Bytes for conversations about Kubernetes, cloud native infrastructure, platform engineering, security, AI infrastructure, and the CNCF ecosystem.

5
out of 5
13 Ratings

About

Kubernetes Bytes is a podcast bringing you the latest from the world of cloud native data management. Hosts Ryan Wallner and Bhavin Shah come to you from Boston, Massachusetts with experienced backgrounds in cloud-native tech. They'll be sharing their thoughts on recent cloud native news and talking to industry experts about their experiences and challenges managing the wealth of data in today's cloud-native ecosystem.

You Might Also Like