Adventures in DevOps

Will Button, Warren Parad

Join us in listening to the experienced experts discuss cutting edge challenges in the world of DevOps. From applying the mindset at your company, to career growth and leadership challenges within engineering teams, and avoiding the common antipatterns. Every episode you'll meet a new industry veteran guest with their own unique story.

  1. לפני 23 שע׳

    Why does anyone use Crossplane?

    Share Episode                    Pushkar Gopalakrishna, Senior Staff Software Engineer at Snap, previously Cruise and AWS, joins to explore why engineering organizations pivot away from Terraform and HCL toward Kubernetes-native tools even when they might not be better. We unpack how developer friction, copy-pasted control structures, and misaligned organizational incentives create massive tech debt—often forcing SREs to manually update infrastructure repositories for compliance, resulting in broken pipelines and severe operational friction.           We debate over the mechanics of Crossplane, detailing how its continuous reconciliation loop and Custom Resource Definitions (CRDs) allow teams to express cloud infrastructure as YAML alongside their application manifests at potentially the cost of async validation. Pushkar pulls back the curtain on how Cruise managed infrastructure at scale using a custom internal platform called 'Juno' to bootstrap GCP projects, repositories, and permissions, while leveraging Crossplane for application-level resources. We also dive into the dangers of using CI tools for continuous deployment, detailing a terrifying incident where a pipeline bug accidentally marked three production Kubernetes namespaces for deletion, and how moving to ArgoCD and Argo Rollouts helped prevent future outages for autonomous vehicles.           Finally, we touch on the realities of non-production environment isolation, testing against live APIs, and why platform teams must balance providing a seamless developer experience without stripping away developer accountability.           💡 Notable Links:          CrossplanePodcast Guest Request for Principal Engineer — What work are you doing?Amazon Multi-level fullyment center for drones✨ Episode: Terraform vs OpenTofu🎯 Picks:          Warren - Books: The Murderbot DiariesPushkar - DJI mini drone

  2. 24 ביולי

    When knowledge is free but the infrastructure isn't

    Share Episode                    As it turns out, the entire artificial intelligence boom is essentially running on Wikipedia's free labor, but while knowledge is free, physical server infrastructure definitely is not. We sit down with Moriel Schottlender, Principal Systems Software Engineer at the Wikimedia Foundation, to dissect how public systems survive an endless onslaught of high-volume AI scrapers and aggressive crawlers. Because 65% of the resource-heavy requests originate from automated bots, we explore how Wikimedia navigates this traffic without blocking legitimate users. We skip the approaches of IP-banning which doesn't work in practice and discuss actual mature architectural strategies, by focusing on the users' needs. From structured database dumps and high-volume enterprise APIs to rate-limiting and CDN caching trade-offs.           It's a mind-bogglingly complex ecosystem ­of open-source, a 25-year-old PHP monolith supporting over 900 distinct site instances across 300 languages and 11 unique projects. It's an immense engineering challenge to modernize infrastructure while serving 250,000 active volunteer editors who build custom workflows via Toolforge—Wikimedia's internal, open-source mini-AWS.           Finally, we have to tackle the philosophical divide between artificial statistical models and human creativity. Because LLMs are trained to predict the statistical mean, they inherently miss the edge cases where real human value, internationalization, and accessibility actually reside. And even if they did, we managed to squeeze out every last bit of AI creativity that early models had until what we are actually left with is the most boring result. We also commiserate over the gratuitous low-quality AI pull requests flooding open-source repositories, drawing parallels to the chaotic Hacktoberfest spam of years past.           💡 Notable Links:          Frodo projectImpact of crawlers on Mediawiki's infrastructureBook: The Platform RevolutionMoriel's LLM experiments✨ Episode: 🎯 Picks:          Warren - Video: Are all flags Drawable in PowerPointMoriel - Audiobook: Dungeon Crawler Carl

  3. 10 ביולי

    Technically We Have Code Reviews and the LLM Semantic Layer

    Share Episode                    We are joined this week by Mark Hay, CTO and co-founder of TextQL and former lead of Text Classification Infrastructure at Meta, to uncover the hidden complexities behind massive-scale machine learning. Mark explains why the most crucial features for identifying abusive behavior, like drug dealers or scammers on Facebook and Instagram, rarely rely on the content itself but instead analyze the underlying behavioral graphs, such as abnormal friend requests or messaging patterns.           Of course we review the adversarial nature of spam detection, where bad actors constantly evolve from simple regex evasion to embedding messages inside images or even utilizing pure symbolic communication, like comparing different sized cucumber emojis to evade text filters. That requires diving into the evolution of database querying and the rise of the semantic layer. Mark unpacks why relying on raw LLMs to write complex SQL is a recipe for hallucinations, and how implementing a "correct by construction" semantic layer guarantees structurally sound queries by restricting outputs to a strictly defined configuration. However, this rigid structure fundamentally stifles the creative flexibility of LLMs.           Lastly, we can't avoid exploring the tension between these approaches and how new tools aim to bridge the gap by dynamically balancing raw SQL generation with structured ontological constraints, providing rapid time-to-value for analytical workflows. Finally, we discuss the controversial philosophical shift occurring within software engineering, particularly the tension between the "Don't Repeat Yourself" principle and "Locality of Behavior".           💡 Notable Links:          ✨ Episode: Semantic Search✨ Episode: Formal Verification✨ Episode: Subjective Model Embeddings🎯 Picks:          Warren - Article: I Left Port 22 Open on the Internet for 54 DaysMark - Hotel Room Exercise: Burpies

  4. 26 ביוני

    Who Needs Testers Anyway?

    Share Episode                              We sit down with Itacama CEO Pia Wiedermayer to discuss the absurdity of siloed QA, the disaster of AI-generated API tests, and why developers hate the word "quality." This time we are asking the age-old question: Who needs testers anyway? Pia and Warren discuss how to dismantle the toxic culture of isolated quality assurance.           We explore how the ghosts of waterfall development still haunt modern teams, creating silos where developers blindly throw unverified code over the wall and expect a separate QA department to magically inject quality. Included is the inevitable discussion on the psychological safety of hiding behind narrow job titles and why refusing to take collective ownership of a product is a guaranteed recipe for architectural failure.           Of course we can't adoiv commenting on the terrifying reality of replacing human intuition with automated hype. Pia shares a case study of a scale-up that aggressively pivoted to "full steam AI development," intentionally excluding both their Product Owner and QA from the entire experiment. Predictably, it did not end well, but we were able to laugh at the painful irony that an AI-accelerated project scheduled for four weeks ended up taking eight weeks, proving that simply generating code without human oversight just creates more sophisticated bottlenecks.           🎯 Picks:          Warren - Wason Selection Task on The Rest Is SciencePia - Book: The Culture Map

  5. 12 ביוני

    What If Tools Are Not Expensive To Build

    Share Episode                    Developers spend more than 50% of their time reading code, making it the single largest expense in software engineering. Despite this massive cost, the industry rarely discusses or optimizes how we read code. So we've brought in Tudor Girba, CEO at Feenk to help us rethink, just how software engineering should be done. Instead of relying on manual reading and generic text editors, teams must shift toward building deterministic, contextual tools to directly extract information and answer questions about their systems.           The suggested solution? Contextual and composable micro-tools writen by everyone focused on exposing just the right information at the right time. This creates the opportunity for structural interrogation of your solution.           And how many tools should we? We'll if one example of tool is testing, and 50% or more of your code can be tests, imagine what percentage of your software should be actually production related!           Most importantly, generic tools fall short, but where can we find how to build the right tools, listen in to find out....           💡 Notable Links:          ✨ Episode: IDE & Copilot & Critical ThinkingBook: Moldable software developmentWardley MapGuest Request: Formal Verification🎯 Picks:          Warren - The real stuff: Underwood Ranches SrirachaTudor - The beaches of Normandy

  6. 5 ביוני

    DR: Staying resilient in the cloud

    Share Episode                    Welcome back to another hopefully, relief from architectural existential dread. This week, we've pulled in Seth Eliot from Arpio, (Ar-Pi-O, RPO, get it?), to dive headfirst into the beautiful, deeply expensive illusion that migrating your legacy infrastructure to a major hyperscaler magically grants it instant immortality. It doesn't. We break down the shared responsibility model for resilience, which was conveniently cribbed straight from the security model, and analyze how the foundational promise of automated fault isolation boundaries routinely crumbles.           From cloud providers sticking multiple "independent" availability zones inside the exact same physical building, to multi-AZ cascading anomalies, to regional power grid failures, it's clear your provider's abstractions aren't nearly as resilient as their marketing slides suggest.           Discussed within is the "Thundering Herd" phenomenon, that can't be ignored even when the failover clusters are designed correctly. From cross-organization KMS re-encryption loops to the horror of fragmented application logs across CloudFront edge regions, at the end of the day, true resilience isn't achieved by forcing your engineering team to implement features, it's about architecting your baseline, confidentiality for the inevitability of production burning to the ground.           💡 Notable Links:          ✨ Episode: Eat your security vegetables✨ Episode: Matt vibecodes✨ Episode: on DNS and isolation🎯 Picks:          Warren - Book: Moldable software developmentSeth - Lockpick set

אודות

Join us in listening to the experienced experts discuss cutting edge challenges in the world of DevOps. From applying the mindset at your company, to career growth and leadership challenges within engineering teams, and avoiding the common antipatterns. Every episode you'll meet a new industry veteran guest with their own unique story.

אולי יעניין אותך