The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering

Fexingo

Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your system is already in production? Lucas and Luna don't pretend to have final answers — they build the conversation so you can draw your own. If you've ever argued about whether a page was necessary or whether an SLO should be tightened, this is your show. #SiteReliabilityEngineering #SRE #Uptime #ProductionEngineering #IncidentResponse #ErrorBudgets #SLOs #Postmortem #ToilAutomation #CapacityPlanning #Observability #DevOps #PlatformEngineering #Resilience #OnCall #FexingoBusiness #BusinessPodcast #Technology Keep every episode free: buymeacoffee.com/fexingo

  1. 1d ago

    How SRE Teams Use Cost-to-Serve Analysis to Optimize Infrastructure

    In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams are applying cost-to-serve analysis to optimize cloud infrastructure spending. Using a real case from a mid-size fintech company that saved $2.3 million annually by mapping compute costs to specific revenue-generating services, they break down what cost-to-serve actually means in an SRE context—allocating shared infrastructure costs to individual products or features. Lucas explains the technical challenges: tagging chaos, shared databases, and the tension between granularity and overhead. Luna pushes back on whether this is just accounting theater or a real operational lever. They discuss practical implementation: starting with a pilot service, using tools like AWS Cost Explorer and Kubernetes cost allocation, and building a culture where engineers see cost data alongside latency and error rates. The episode also covers common pitfalls like over-aggregation and the risk of optimizing costs at the expense of reliability. Listeners will walk away with a concrete framework for starting cost-to-serve analysis in their own teams. #CostToServe #CloudCostOptimization #FinOps #SRE #SiteReliabilityEngineering #InfrastructureCost #AWS #Kubernetes #CostAllocation #CloudEconomics #Technology #FexingoBusiness #BusinessPodcast #Podcast #DevOps #Observability #CostEfficiency #EngineeringCulture Keep every episode free: buymeacoffee.com/fexingo

    How SRE Teams Use Cost-to-Serve Analysis to Optimize Infrastructure
  2. 3d ago

    How SRE Teams Use Capacity Planning to Prevent Outages Before They Happen

    Lucas and Luna dive into the science of capacity planning for site reliability engineering. They break down how Netflix uses predictive modeling to scale infrastructure ahead of demand spikes, avoiding the kind of cascading failures that hit other streaming services during major events. The episode explores real-world examples of capacity planning failures—like the AWS outage that took down half the internet in 2021—and explains the difference between reactive scaling and proactive capacity planning. Lucas shares a concrete framework for setting capacity thresholds using the '80/20 rule' and discusses how SRE teams at companies like Google and Uber use historical traffic patterns to forecast resource needs. Luna challenges the assumption that more capacity is always better, pointing to the 2024 Google Cloud outage caused by overprovisioning. The hosts also touch on the role of automation in capacity planning and the importance of regular load testing. If you've ever wondered why some services stay up during traffic surges while others crash, this episode explains the engineering behind the scenes. #CapacityPlanning #SRE #SiteReliabilityEngineering #Netflix #AWS #GoogleCloud #Uber #PredictiveModeling #InfrastructureScaling #OutagePrevention #LoadTesting #Automation #Tech #Engineering #FexingoBusiness #BusinessPodcast #Technology #SREPodcast Keep every episode free: buymeacoffee.com/fexingo

    How SRE Teams Use Capacity Planning to Prevent Outages Before They Happen

About

Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your system is already in production? Lucas and Luna don't pretend to have final answers — they build the conversation so you can draw your own. If you've ever argued about whether a page was necessary or whether an SLO should be tightened, this is your show. #SiteReliabilityEngineering #SRE #Uptime #ProductionEngineering #IncidentResponse #ErrorBudgets #SLOs #Postmortem #ToilAutomation #CapacityPlanning #Observability #DevOps #PlatformEngineering #Resilience #OnCall #FexingoBusiness #BusinessPodcast #Technology Keep every episode free: buymeacoffee.com/fexingo

You Might Also Like