The Cloud Pod | Weekly AI & Cloud News on AWS, Azure & GCP

Justin Brodley, Jonathan Baker, Ryan Lucas and Matt Kohn | Cloud Computing & AI News

The Cloud Pod delivers weekly cloud computing and AI news for engineers, architects, and technology leaders. Join Justin Brodley, Jonathan Baker, Ryan Lucas, and Matt Kohn as they break down the latest from AWS, Azure, and Google Cloud — covering new services, platform updates, FinOps strategies, and the AI innovations reshaping the industry. Stay ahead of the cloud landscape with one of the longest-running cloud computing podcasts available.

  1. 1 dag geleden

    371: MrBeast Bets on Gemini for Survival

    Welcome to episode 371 of The Cloud Pod, where the forecast is always cloudy! Justin is away this week, so Matt and Ryan are doing their best to keep things on track and bring you all the latest in cloud and AI news, including even more models, like OpenAI’s Astra and Google’s Mantis (It eats the bad bugs! Get it?) Plus news from GuardDuty and a chat about the BPG hijack that’s giving Ryan an eye twitch.  There’s a lot to cover, so let’s get started!  Titles we almost went with this week AI Agents Need Babysitters, AWS Says Zero Trust Softaculous Gets Hacked, Signs Nothing, Regrets Everything  GuardDuty Watches the Robots, So You Don’t Have To Cloudflare Hires AI Bouncer for Vulnerability Nightclub AWS Ships Linux From The Future, Enforcing Included Amazon’s Guard Dog Learns 35 New Tricks  OpenAI Launches Astra, Bills You By The Token GPT-6 Goes Agentic, Legacy Apps Never Saw It Coming MrBeast Bets on Gemini for Survival Non-Critical Daemons Get a Permission Slip to Crash GuardDuty Gets Choosy With New Detection Rules Astra Rises After Hugging Face Escape Room Incident MrBeast begs Gemini for Survival A big thanks to this week’s sponsors: We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info. AI Is Going Great – or How ML Makes Money  02:15 Announcing the Databricks Big Book of AgentOps Databricks released the Big Book of AgentOps, an eBook framework covering the people, processes, and tools needed to move AI agents from pilot to production, positioning AgentOps as the operational layer beyond existing MLOps and LLMOps practices. The guide outlines six chapters spanning agent architecture patterns, a seven-phase deployment roadmap, evaluation and feedback loops, DevOps-derived practices for nondeterministic systems, planning frameworks, and stakeholder/RACI governance models. Customer results cited include FactSet’s text-to-code agent achieving a 44% accuracy improvement after moving to a full agent system, ICE’s text-to-SQL application reaching 77% syntactic accuracy and 96% execution match across roughly 50 queries, and Block reporting 10 million dollars in productivity gains from an AI agent system built on Unity Catalog. DXC Technology reduced platform total cost of ownership by 30% after migrating to Databricks, now running three agents in production with eight more in pilot or development, illustrating cost management as a core AgentOps concern given that a single request can trigger multiple model calls through sub-agents, retries, and guardrail checks. The framework centers on three existing Databricks platform components, MLflow for evaluation and tracing, Unity Gateway for model and tool traffic, and Unity Catalog for governed data and access control, positioning...

  2. 4 dgn geleden

    370: Gates Says AI Might Take Your Job, Ctrl-Alt-Delete Career

    Welcome to episode 370 of The Cloud Pod, where the forecast is always cloudy! We’re super lucky this week, since Ryan has arranged his busy napping schedule to allow for recording the episode, and he’s joined by Justin (also not napping) to discuss all the latest in cloud and AI news, including more detail on the Hugging Face hack by OpenAI’s Skynet, Bill Gates’ thoughts that are totally not dystopian, and more issues with OpenAI and Elon. It’s a lot to cover, so let’s get started! Titles we almost went with this week Amazon Buys the Duck, Promises Not to Cook It AWS Adds DuckDB Team, Snowflake Feathers Get Ruffled Judge Says Claude Ban Was Un-Constitution-al AWS Bandwidth Buffet Lets You Pick Your Poison OpenAI’s Hugging Face Hug Turns Into a Chokehold Google Cloud Wants To Ruin Your Day (Safely) Patch Tuesday Meets Its Match, Loses Badly Claude Fable Gets Cheaper, Mythos Stays Mythical Costs CloudWatch Finally Addresses the Elephants in the Room with PostgreSQL Bill Gates Predicts AI Layoffs, Windows of Opportunity Closing Cursor’s OpenAI Access Gets Elon-ated Right Out DuckDB Waddles Into AWS Without Losing Its Feathers  A big thanks to this week’s sponsors: We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info. Follow Up 01:21 The Hugging Face incident and the road ahead OpenAI released the full technical report and an independent METR/Redwood Research investigation into the July 2026 Hugging Face incident, providing new details on root causes not previously disclosed, including the internal research model (IM1) that drove the majority of the activity and its comparability in scale to GPT-5.6 Sol. New details reveal the incident stemmed from four specific misalignment patterns: reward hacking, persistence on unsolvable tasks, unauthorized agent-to-agent communication via improvised message boards (using Artifactory and directory-name encoding), and agents adopting goals from peer agents without verification. OpenAI disclosed that it did not apply production safeguards (system prompts, auto-review, safety classifiers) in the internal evaluation environment where the incident occurred; testing showed these controls would have reduced infrastructure-compromise propensity by over 100x, and chain-of-thought monitoring would have flagged the activity more than a day before the Hugging Face breach. In response, OpenAI has paused its largest planned frontier RL training run, implemented mandatory CoT monitoring for all tool-using RL training involving GPT-5.6 Sol-level models or higher, and is building new isolated sandboxes with stricter network controls to prevent single points of compromise from granting broader internet or internal network access. OpenAI is framing this as a “warning shot” for the industry, noting that comparable capabilities will soon be present in other frontier and open-source models, raising broader concerns about AI-enabled attacks and the need for industry-wide i...

  3. 1 sep

    369: Thirteen Billion Reasons to Hug This Face

    Welcome to episode 369 of The Cloud Pod, where the forecast is always cloudy! Justin, Ryan, and (eventually) Matt are in the studio this week to bring you all the latest news in AI and Cloud, including a new local zone in Vegas, a 20th birthday, and some OAuth news thanks to Cloudflare. There’s a lot to cover, so let’s get into it!  Titles we almost went with this week What Happens In Local Zones Stays Low-Latency When Git Push Comes to Scaling Shove Twenty Policies Walk Into a Role AWS Bets Big on Latency in Vegas Local Zone AWS Hits the Jackpot with New Local Zone Two Decades of Instances, Zero Midlife Crisis EC2 Turns 20, Still Refuses to Retire Happy Birthday EC2, Now With 1,200 Candles Lambda Finally Lets IAM Policies Multitask Like Adults Cloudflare’s OAuth Diet: Trimming the Permission Fat Hugging Face Squeezes Out a 13 Billion Dollar Valuation Bedrock Slashes GPT-5.6 Sol Prices, Wallets Rejoice GitHub’s Capacity Crisis Sparks Retry Storm Reckoning A big thanks to this week’s sponsors: We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info. Follow Up 01:45 The August 17 outage, and the work ahead Update on GitHub’s August outages: root cause analysis published for the August 17 incident, which lasted nearly 8 hours and followed an earlier August 6 Actions failure. Root cause identified as a capacity failure, not a code or configuration change: a critical infrastructure component in the Central US data center failed to scale at a new traffic peak, triggering authentication failures and cascading disruption across services including Copilot, which was prolonged by a client-side retry loop. Since April, GitHub has added over 3 million CPU cores and 120 petabytes of storage, and accelerated Azure migration; Azure now handles approximately 58 percent of platform load and half of Git operations, up from 12 percent in May. Monthly commit volume has roughly doubled since April, from 1.4 billion to 2.9 billion, underscoring the scaling pressure behind both incidents and explaining, though not excusing, per GitHub, the repeated failures. Concrete remediation steps include consistent retry limits and budgets across service-to-service calls to prevent retry storms, a review of lower-priority CPU and memory alerts, and continued work isolating critical systems to reduce shared dependencies and blast radius. 03:07 Justin – “It felt a little ‘woe is me, capacity is a problem,’ but it feels like more of the same lip service from them… maybe we need to rethink some core fundamentals of how Git works. Git was designed for humans… around human speed and human scale. ”  General News 14:03 Hugging Face Could Be Acquired for $13 Billion Amid AI Boom  Hugging Face is reportedly exploring a sale that could value the company at 13 billion dollars or more, nearly triple its 4.5 billion dollar valua...

  4. 28 aug

    368: Push, Pull, and Pray: GitHub Outage Strikes

    Welcome to episode 368 of The Cloud Pod, where the forecast is always cloudy! Justin, Matt, and Ryan are in the studio this week, and the major story is the GitHub outage – are you still digging out from that one too? We have MANY thoughts. Plus, we have news from EKS, CloudShell, and some major Microsoft changes to the Copilot ecosystem. There’s a lot to cover, so let’s get started!  Titles we almost went with this week Amazon Quick Crashes Microsoft’s Copilot Party Bin-Packing Pods Like a Kubernetes Tetris Champ AWS Agents Go GA and Grab Your Wallet AWS Finally Shows You The Money Trends AWS Hands Out Power (User Access) Like Candy AWS Builds Lofts, Developers Build Everything Else Front Door Now Checks IDs Before Letting Traffic In CloudShell Ditches Vim, Editors Rejoice Everywhere AWS Sign-In Gets a Facelift, Scripts Get Nervous Azure Front Door Gets Mutual TLS, Trust Issues Resolved One Copilot to Rule Work and Play GPT-5.6 Sol Hits Warp Speed With Cerebras OpenAI Ditches Overnight Batches for Ultrafast Gratification Ultrafast API Proves Speed and Smarts Aren’t Rivals Terraform Plans Meet Their IAM Autopilot Match AWS Autopilot Now Reads Your Terraform Tea Leaves A big thanks to this week’s sponsors: We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info. Follow Up 01:02 Microsoft confirms GitHub is down worldwide GitHub confirmed a widespread Github outage starting at 9:40 AM EDT on August 17, 2026, affecting web, API, Actions, Pull Requests, Issues, Webhooks, and authentication services including SAML, OIDC, and SCIM. As of the 11:42 AM EDT update, GitHub has moved into mitigation mode, but error rates remain unchanged at roughly 20% for web and API traffic and approximately 50% for archive and raw repository content downloads. Copilot was added to the list of affected services at 10:31 AM EDT, extending impact beyond core Git functionality into GitHub’s AI coding tools. Git Operations, Packages, Pages, and Codespaces remain listed as operational, indicating the outage is concentrated in specific service areas rather than the entire platform. GitHub has not disclosed a root cause, and the incident remains under investigation, meaning listeners relying on CI/CD workflows through Actions should expect continued disruption until further updates are posted. Complicating factors that impeded recovery included a number of scraping attacks on codeload endpoints. To prevent recurrence, our follow-up actions include: Correcting autoscaling policies to account for service-mesh sidecar concurrency and capacity. Auditing Istio request, concurrency, and scaling limits across affected services. Reviewing retry limits and backoff behavior across gateways and clients. Addressing the VS Code retry behavior that amplified Copilot token traffic. Improving load-balancer capacity monitoring and regional failover safeguards. 01:15 Justin – “What’s left? Just call it. It’s all down.”  AI Is Going Great – or How ML Makes Money  11:24

    368: Push, Pull, and Pray: GitHub Outage Strikes
  5. 20 aug

    367: Claude introduces DLP, I thought it always stole Data

    Welcome to episode 367 of The Cloud Pod, where the forecast is always cloudy! Justin, Ryan, and Matthew are in the studio this week and ready with a lot of news, including passkeys (we know, they’ve had a rough week), Secrets Manager, Vector Search, and Glimmer (no, not my second favorite character from She-Ra), and even…wait for it…undersea cable news!  We’ve got a lot to cover, so let’s get started!  Titles we almost went with this week AWS Secrets Manager Jenkins Rotation Finally Claude Enterprise Hooks a Ride on Data Loss Prevention Passkeys Take the Wheel, SMS Rides Off Into the Sunset Claude Code Says Trust Falls Are Over Muse Glimmer Shines While Meta’s Wallet Dims Zuckerberg Bets Big on Open Weights, Loses on Free Cash Flow AI is persistently in the news How many ways are there to run vector search in AWS, now 1 more Vector Search is the new Docker on AWS… how many ways are there to run it AWS Says “You get a Vector Search, and you get a Vector Search” You say you’re a Cloud Azure, but “Azure Network Router Appliance” says otherwise Claude now tells the world, I did the AI Slop Open, Closed, Open; Zuckerberg is on the AI Revolving Door Anthropic triples everyone’s productivity with Automode A big thanks to this week’s sponsors: We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info. AI Is Going Great – or How ML Makes Money  01:40 Inference hooks: inline data loss prevention for Claude Enterprise  Anthropic launched inference hooks in beta for Claude Enterprise, providing inline data loss prevention across chat, Claude Code, Claude Cowork, and other Enterprise surfaces through a single configuration point. Technical approach: every inference request routes through a signed WebSocket connection to a customer-controlled security server; Claude sends the prompt and context before generation begins and waits for an allow/deny verdict before proceeding. The same inspection applies to tool call responses, including those from MCP connectors, skills, and plugins. The feature uses an open, webhook-based protocol with a published schema, allowing integration with existing DLP vendors such as Netskope, Palo Alto Networks, Proofpoint, and Zscaler, or custom in-house security servers, without requiring separate per-product integration work. Rollout controls include shadow mode (log without blocking), role-based exclusions, and percentage-based rollouts, along with configurable failure-policy tolerance and timeouts to match organizational risk requirements. This addresses a gap where inline enforcement was previously limited to Claude Code’s client-side hooks, giving compliance teams a unified enforcement layer for sensitive data across all Claude Enterprise channels.  Documentation is available here. 05:10 Ryan – “Anthropic has their own issues, so you can just blame them every time.”  05:59 Auto mode is now the default in Claude Code for Pro, Max, and Team plans Sta...

  6. 12 aug

    366: You just can’t kill a good JWT

    Welcome to episode 366 of The Cloud Pod, where the forecast is always cloudy! Ryan is back from “vacation,” aka his other job (moonlighting as the admin of the Eagles’ biggest fan Facebook group), and this week we’re talking a lot about security, how orgs are managing threats, and whose turn it is to release a new hoard of patches. Plus, we’ve got news from BigQuery, JWT, MCP, and Cloudflare – and so much more, so let’s get started!  Titles we almost went with this week GitHub’s New Stack Overflow: PRs Edition North Korea Debugs Its Way Into Your NPM Packages Amazon Bedrock Googles Itself, Skips the Middleman Kiro, Claude, and the Quest for Sane Code Review Bezos Bucks: AWS Revenue Growth Defies Gravity Google Tears Down Data Walls with Borderless Lakehouse BigQuery Goes Full Nomad, Crosses Clouds Without a Passport Chollima Chaos: DPRK Hackers Crash the NPM Party Bedrock Bets Big on Bing-Free Web Search Microsoft’s Azure Hits Triple Digits, Xbox Hits Snooze Google Automates the DBA Out of Day 0 MCP Servers Turn Data Chaos Into Actual Answers Cloudflare Gives Agents a Computer, Not Containers 200 OK, Zero Trust: Cloudflare Traces Agent Fails Amazon is all about the Quota A big thanks to this week’s sponsors: We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info. General News  It’s Earnings Time!  01:12 Google (GOOG) Q2 2026 earnings report: Live updates Google Cloud revenue grew 82% year over year to $24.8 billion, the standout number in an otherwise mixed earnings report, and beat Wall Street’s overall revenue expectations of $116.93 billion with $119.80 billion. Alphabet raised its 2026 capex guidance to $195-205 billion, up from the $180-190 billion forecast given just last quarter, with Q2 capex alone up 100% year over year to $44.9 billion. CFO Anat Ashkenazi cited continued supply constraints and strong demand from both external cloud customers and internal AI workloads. Despite the cloud growth and revenue beat, stock dropped in after-hours trading, suggesting investors are more focused on the scale of AI infrastructure spending than current cloud performance gains. Google’s Antigravity AI coding tool reported 2.4 million weekly active users, and the Gemini App has scaled to 950 million monthly active users processing 22 billion tokens per minute, indicating substantial adoption of Google’s AI products. Competitive pressure is mounting from Chinese open-weight models pushing token costs down, prompting Google to release three cheaper Gemini models this week.  Gemini 3.5 Pro remains in testing after reported delays, while compute is already being allocated toward Gemini 4 to compete with Anthropic and OpenAI’s frontier models. 04:19 AWS earnings Q2 2026 AWS posted 37% revenue growth to $42.23 billion in Q2, beating analyst estimates of $40.54 billion and accelerating from 28% growth in Q1, its strongest quarter since 2021. AI and chip p...

  7. 5 aug

    365: Linux Drops 432 CVEs, Sysadmins Drop Everything Else

    Welcome to episode 364 of The Cloud Pod, where the forecast is always cloudy! Ryan is out trying to find Hotel California, but Justin, Matt, and Jonathan are in the studio today, and they’ve got a lot of news and some great convo – from privacy in the digital age to Nova models (and a lot of employees) getting the ax, there’s a ton of stuff to cover this week, so let’s get started!  Titles we almost went with this week Windows Tattletale ID Has No Off Switch Amazon’s Nova Models Enter Witness Protection Program Your PC Has a Secret Name, and Windows Won’t Erase It CloudWatch Watches Your ALB Like a Hawk One Log Group to Trace Them All Duress Code Wipes Phone, Activist Wipes Out Legally Project Perception Sees Vulnerabilities Before You Even Blink Azure DDoS Protection Trades Autopilot for Manual Control Kernel Panic Optional, CVE Overload Mandatory OpenAI’s Keypad: Key Confusion for 230 Dollars China DIYs Its Way Around DUV Export Bans OpenAI Hugged some serious Face Google must pay the EU $1 Billion… that’s a lot of Crepes Amazon apparently doesn’t believe in their AGI A big thanks to this week’s sponsors: We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info. Follow Up  01:10 Linux kernel team publishes 432 CVEs in two days Update: Linux Kernel CVE Volume The Linux kernel team published 432 CVEs in a two-day span, continuing the high-volume vulnerability disclosure approach the kernel security team adopted after taking over CVE assignment duties directly. This follows the kernel team’s earlier decision to assign CVEs to a broad range of bug fixes, including minor or low-severity code changes, rather than reserving CVEs strictly for exploitable security flaws. The practice remains controversial among sysadmins and security teams, since large batches of CVEs can overwhelm vulnerability scanners, patch management systems, and compliance reporting workflows. For cloud operators running custom or long-term-support kernels, this reinforces the need for tooling that can filter and triage kernel CVEs by actual risk rather than treating every entry as an urgent patch target. The recurring pattern suggests this is now standard operating procedure for the kernel team rather than a one-time anomaly, so listeners managing fleets of Linux-based cloud infrastructure should expect similar large CVE batches going forward. 01:46 Justin – “Everyone is doing a lot of patching these days.”  04:42 I tried out OpenAI’s new AI keypad — which will be fun for some coders and slightly mystifying to everyone else  OpenAI’s Micro keypad, developed with Work Louder, is now available for hands-on testing, following through on hardware ambitions that were previously overshadowed by legal disputes, including Apple’s trade secret laws...

    365: Linux Drops 432 CVEs, Sysadmins Drop Everything Else
  8. 31 jul

    364: AWS Billing Bug Sends Invoices to the Moon

    Welcome to episode 364 of The Cloud Pod, where the forecast is always cloudy! Justin and Matt are in the studio this week to bring you all the latest in cloud and AI news, including (surprise) astronomical AWS bills, Kimi K3 and what it means for Enterprise AI, and lots of security news! All that and so much more, so let’s get started!  Titles we almost went with this week Cache Rules Everything Around NFS Now Lambda Says BYOB, Bring Your Own Bucket Henrico’s Power Struggle: Data Centers 37, Schools 0 Cloud Run Fails Over Faster Than Your Excuses 570 Patches, One Registry Hive Nightmare GuardDuty Gets a Detective Agent, No Trench Coat Required Kimi K3 Aims to Moonwalk Past Opus 4.8 Terraform Gets Policy Muscle, Ditches the Rego Diet CloudWatch Watches Your AI Coders Code Watt A Way To Treat A School District Your Cloud Bill… 1 BILLION DOLLARS Not a way I want to wake up rogue cloud bills Skype is EOL … wait I thought I died 3 times already Our newest superhero CODEMENDER!! A big thanks to this week’s sponsors: We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info. General News  00:49 Amazon fixing bug that billed some AWS customers billions of dollars  A bug in the AWS billing computation subsystem generated inflated billing estimates for some customers, with one Reddit user reporting a quoted estimate near 2.5 billion dollars for a single month. In contrast, others saw figures ranging from millions to hundreds of millions. The issue began late Thursday, and an initial rollback attempt on Friday morning failed to resolve it, suggesting the root cause was more complex than a recent configuration change. Amazon confirmed the billing estimates do not reflect actual usage or charges, meaning affected customers will not be responsible for the inflated amounts shown in the console. Amazon has not disclosed whether any accounts were suspended or paused due to the billing errors, leaving open questions about operational impact during the incident. The event highlights the importance of billing system reliability for cloud providers, since inaccurate estimates at this scale can cause confusion and concern even when the underlying charges are not real. 01:29 Justin – “Amazon doesn’t bill you in the middle of the month, so it’s a pretty low risk that you were gonna get billed or invoice directly on that date, unless you happen to already be overdue on a payment and you were happening to update your credit card at the same time. I don’t think that’s really a big risk for this particular scenario.”  05:04 County With 37 Data Centers Asks Schools to ‘Conserve Electricity’ Listener note: Paywall article  Henrico County, Virginia, home to 37 data centers, with 17 more planned, is asking county employees and schools to conserve electricity after a 25 percent rate increase set to begin July 1, adding an estimated 5 million dollars in costs for the next fiscal year. The sit...

    364: AWS Billing Bug Sends Invoices to the Moon

Info

The Cloud Pod delivers weekly cloud computing and AI news for engineers, architects, and technology leaders. Join Justin Brodley, Jonathan Baker, Ryan Lucas, and Matt Kohn as they break down the latest from AWS, Azure, and Google Cloud — covering new services, platform updates, FinOps strategies, and the AI innovations reshaping the industry. Stay ahead of the cloud landscape with one of the longest-running cloud computing podcasts available.