The Data Flowcast: Mastering Apache Airflow ® for Data Engineering and AI

Astronomer

Welcome to The Data Flowcast: Mastering Apache Airflow ® for Data Engineering and AI— the podcast where we keep you up to date with insights and ideas propelling the Airflow community forward. Join us each week, as we explore the current state, future and potential of Airflow with leading thinkers in the community, and discover how best to leverage this workflow management system to meet the ever-evolving needs of data engineering and AI ecosystems. Podcast Webpage: https://www.astronomer.io/podcast/

  1. 4d ago

    Building self-healing Airflow pipelines at ATC Drivetrain

    Manufacturing data pipelines can't afford silent failures. When quality decisions and shop-floor visibility depend on Airflow, a broken DAG at midnight can cost real money. In this episode, Kenten Danas talks with Kumuda Sreenivasa, Founder of Receitly and Senior Data Architect at ATC Drivetrain, about how her team built self-healing Airflow pipelines, where AI fits into the recovery loop, and how they orchestrate AI agents as governed workflow components. Key Takeaways: 00:00 Introduction. 02:29 How ATC Drivetrain uses Airflow across thousands of DAGs to orchestrate ETL/ELT jobs, data quality checks, and production reporting for a complex automotive remanufacturing environment. 04:08 Defining self-healing: pipelines that identify a known failure, decide whether they can recover safely, execute an approved action, and validate the result, all without paging an engineer. 07:03 The five-layer self-healing architecture: observe, classify, policy, recover, and validate. 10:00 Where AI fits in the recovery loop: classification, context gathering, and recommendations, but never bypassing operational policies or approval steps. 13:00 A concrete before/after: a currency-exchange failure caught overnight by AI-assisted recovery that saved four hours of downtime and roughly 120K. 14:27 Confidence levels and success rates: about 95% of small failure modes recover on their own. 15:34 AI-assisted troubleshooting at scale: how contextual log analysis and recommended actions save engineers from digging through thousands of log lines. 20:13 Orchestrating AI agents through Airflow: treating agents as bounded, governed workflow components with human approval and confidence-based stop conditions. 23:45 What Kumuda wants next from Airflow: stronger AI agent governance, standardized tracking of prompts and tool calls, and more flexible event-driven execution. Resources Mentioned: Orchestrate EverythingApache AirflowATC Drivetrain Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #Airflow

    Building self-healing Airflow pipelines at ATC Drivetrain
  2. Sep 3

    Orchestrating strictly sequential ETL pipelines at Synechron

    Strict sequential execution across DAGs sounds simple until you have scheduled pipelines and event-driven pipelines writing to the same MongoDB collections. Ivana Isailovic, Senior Big Data Engineer at Synechron, joins Marc Lamberti to walk through the three-layer DAG architecture her team built to solve exactly that, plus how they generate 200-task DAGs and how they rebuilt subdag-style group retries in Airflow 3. Key Takeaways: (00:00) Introduction. (02:16) The stack: Snowflake source, MongoDB with a medallion (bronze, silver, gold) layout, Spark for processing, Airflow for orchestration, Elasticsearch for reports. (05:29) Why standard Airflow options (max active runs, pools, dependency setups) each solved only part of the problem. (06:40) Data consistency across bronze, silver, and gold layers is what forced strict sequential execution. (09:11) Scheduled DAGs versus event-driven DAGs triggered at any moment from the application side. (10:04) The three-layer architecture: trigger DAGs, a single proxy DAG that controls the queue, and main ETL DAGs. (12:00) The queue is literally another DAG. The proxy DAG allows only one active run and serializes everything behind it. (15:13) 200-task DAGs generated from nested task groups and YAML configuration files, with DAG versions tied to release numbers. (17:38) How the layers talk to each other: sensors and TriggerDagRunOperator. (19:44) Migrating from subdags to task groups without losing the ability to retry a whole group. (21:49) Airflow 2.9 approach: reset task instance state via the metadata DB, keyed off the task group identifier. (23:13) Airflow 3 approach: move the retry logic onto the official REST API for stability, security, and maintainability. Resources Mentioned: Orchestrate EverythingApache AirflowSnowflakeMongoDBApache SparkElasticsearch Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #Airflow

    Orchestrating strictly sequential ETL pipelines at Synechron
  3. Aug 27

    Managing financial datasets with Airflow at Wise

    Financial data pipelines have to be right the first time. On this episode, Kenten sits down with [Antonello Benedetto](linkedin.com/in/anbento4), Staff Data Engineer at Wise, to talk about how the central data and analytics engineering team runs Airflow for critical financial datasets, tiers pipelines by reliability, and orchestrates LLM-enabled workflows with validation layers and agent-checking-agent patterns. Key Takeaways: 00:00 Introduction.01:47 What Wise does and Antonello's role in the central data and analytics engineering team, a hybrid platform-plus-analytics team that owns dbt infrastructure, BI, and Analytics MCPs as a service.05:53 Three principles that guide Airflow pipeline design at Wise: a clean separation between orchestration and computation logic, computational awareness (offloading memory-intensive tasks to EMR or SageMaker), and standardized deployments.07:21 Why Wise treats Airflow as a pure orchestration layer and pushes memory-intensive work to external workers.08:45 Moving to the Python Virtual Environment Operator to standardize Airflow deployments across the org while giving analysts and data scientists per-job Python environments.10:50 The tiering system for pipelines, how it distinguishes highly controlled, well-documented, well-observed pipelines from newer ones, and how requirements from downstream drive tier promotion.17:18 Where LLM-enabled workflows differ from standard pipelines: validation layers for specific use cases, plus observability and evaluation platforms that track model performance across executions.19:08 Using Airflow to orchestrate LLM generation of monthly variance commentary for analysts.21:10 Handling non-idempotent LLM outputs with multi-layer validation against source-of-truth data, and using a second agent (CI/CD style) to validate the first agent's output.23:04 How AI-enabled workflow orchestration differs from batch ETL, and why teams should start small before building fully agentic pipelines.25:25 What Antonello would most like to see from Airflow next: native support for agentic workflows and better local development that mirrors production. Resources Mentioned: [Orchestrate Everything](https://astronomer.link/data-flowcast-oe)[Wise](wise.com)[Wise Careers](wise.jobs)[Apache Airflow](airflow.apache.org)[dbt](getdbt.com)[Python Virtual Environment Operator](airflow.apache.org/docs/apache-airflow/stable/core-concepts/operators.html) Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #Airflow

    Managing financial datasets with Airflow at Wise
  4. Aug 20

    Managing travel platform data with Airflow at Headout

    Running a travel platform means dealing with fast-moving inventory, real-time fraud detection, and heavy performance marketing attribution, all while keeping data trustworthy for a data-driven org. In this episode, Mrinalini Singh, Data Platform Engineer at [Headout](headout.com), walks through how her team uses Airflow as the nervous system of their stack: orchestrating dbt with a write-audit-publish pattern, running ML training and inference, and wiring up alerting that points to the exact commit that broke a DAG. Key Takeaways: 00:00 Introduction.01:05 What a data platform engineer does at Headout, and the hub-and-spokes model where analysts and scientists write their own dbt models.04:11 The specific data challenges of a travel platform: fast-changing inventory, real-time fraud analytics, and performance marketing attribution.05:40 Where Airflow sits in the stack, from ingestion to transformation to serving.07:01 The write-audit-publish dbt pattern and why slightly stale data beats wrong data.09:55 Why Headout uses a custom Python operator instead of the dbt provider or Cosmos, reading the dbt manifest to build task groups per model.13:45 ML use cases on Airflow: Feast feature store, model training, inference, and data/feature drift tracking.17:00 Custom Slack failure hooks that stitch together Airflow logs, GitHub commit URLs, and teammate Slack IDs.20:04 A zombie task incident that filled the metadata DB, caused locking issues, and drove the move to Grafana-based monitoring.22:26 Using AI to generate Airflow code, encoding internal patterns as a skill file, and running an AI reviewer bot on every PR. Resources Mentioned: [Orchestrate Everything](https://astronomer.link/data-flowcast-oe)[Headout](headout.com)[dbt](getdbt.com)[Cosmos](github.com/astronomer/astronomer-cosmos)[Feast](feast.dev)[Apache Flink](flink.apache.org)[Grafana](grafana.com) Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #Airflow

    Managing travel platform data with Airflow at Headout
  5. Aug 13

    How Airflow orchestration decisions impact Spark performance

    Airflow and Spark are one of the most common combinations in modern data platforms, but the orchestration decisions made on the Airflow side often determine whether Spark jobs run fast and cheap or slow and expensive. In this episode, Kenten is joined by [Meni Shmueli](linkedin.com/in/meni-shmueli-dataflint), Co-Founder and CEO at [DataFlint](dataflint.io), to dig into how Airflow and Spark fit together, where teams go wrong, and how AI is changing the way they reason about cost and performance. Key Takeaways: 00:00 Introduction.01:40 Meni's background as a data engineer and what led him to start DataFlint, a production observability agent for Apache Spark.02:53 Why the Airflow plus Spark combination is so common, and how each tool plays to its strengths.04:23 How orchestration decisions in Airflow directly impact Spark performance and cost.04:44 A customer story where parallelizing Airflow tasks made Spark jobs slower, less stable, and more expensive, and the opposite case where sequential runs left compute on the table.07:02 The number one mistake teams make benchmarking pipelines: only looking at the Airflow side and ignoring underlying Spark cost and resource usage.08:13 What DataFlint's Airflow and Astro integration gives teams, and how it brings production context into AI agents.10:46 How AI is changing pipeline optimization, including holistic scheduling across hundreds of pipelines and connecting context from Spark, FinOps, and cloud.12:21 A customer migration from Databricks to EMR that cut workflow costs by 80%, with examples of up to 100x optimizations.14:23 Where the Airflow and Spark story could be better, including Spark Declarative Pipelines and tighter feedback between the two projects. Resources Mentioned: [Orchestrate Everything](https://astronomer.link/data-flowcast-oe)[DataFlint](dataflint.io)[Apache Airflow](airflow.apache.org)[Apache Spark](spark.apache.org)[Astro](astronomer.io/astro)[Airflow Summit](airflowsummit.org) Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #ApacheAirflow

    How Airflow orchestration decisions impact Spark performance
  6. Aug 6

    Using Airflow for diverse client projects at Accion Labs

    When a real-time loan eligibility scoring pipeline is built on cron jobs, midnight pages are inevitable. In this episode, [Chandan Gowda](linkedin.com/in/chandan-gowda-a-h-744908195), Data Engineer at [Accion Labs](accionlabs.com), joins Kenten to discuss how his team uses Airflow across client projects, including a financial services scoring use case and a POC applying production-grade orchestration to RAG and GenAI data pipelines. Key Takeaways: 00:00 Introduction.01:00 What Accion Labs does as a technology consulting and services firm working across BFSI, healthcare, and retail.02:00 Chandan's role at the intersection of data engineering and GenAI, building pipelines one week and RAG-based agents the next.04:20 Why Airflow tends to win client evaluations: infrastructure agnostic, no cloud lock-in, fine-grained control over pipeline logic.06:00 Containerizing Airflow on Kubernetes or VMs so migrations between clouds don't require a rewrite.08:14 The loan eligibility scoring use case for a financial services client, and replacing fragile cron jobs with a single Airflow DAG end to end, cutting effort by about 25%.10:42 Triggering strategy: S3 file sensors as the primary trigger handling 90% of runs, plus a scheduled fallback as a safety net.12:53 End-to-end flow inside Airflow: ingestion into the data lake, validation and transformation with credit bureau joins, containerized model inference, and writeback to the loan management system.15:56 The AI orchestration POC and why the data feeding GenAI models needs the same rigor as any production pipeline.18:09 Using Airflow to detect document changes and re-chunk and re-embed only what changed, with quality thresholds and rollback before promoting to live.20:23 The roadmap: model evaluation pipelines, multi-agent orchestration, and provider packages for LangChain, OpenAI, and Hugging Face.22:32 Wishlist for Airflow: native event-driven triggers beyond polling sensors, first-class observability for AI workloads, and better dynamic DAG generation at scale. Resources Mentioned: [Orchestrate Everything](https://astronomer.link/data-flowcast-oe)[Accion Labs](accionlabs.com)[Apache Airflow](airflow.apache.org)[Airflow LangChain provider](airflow.apache.org/docs/apache-airflow-providers-langchain/stable/index.html)[Airflow OpenAI provider](airflow.apache.org/docs/apache-airflow-providers-openai/stable/index.html) Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #Airflow

    Using Airflow for diverse client projects at Accion Labs
  7. Jul 30

    Consolidating legacy job schedulers with Airflow at Trading Technologies

    Migrating two decades of legacy job orchestration is the kind of project most teams quietly avoid. In this episode, Marc Lamberti is joined by [Sanket Patel](linkedin.com/in/sanket-patel-a2a57b3a), Director of Engineering at Trading Technologies, to talk through how his team is consolidating 20 years of C#/.NET-based scheduling and a recently acquired company's stack onto Airflow, Astronomer, and Snowflake. The conversation covers orchestrator selection, observability, the operational economics of managed Airflow, the human side of migrations, and where AI fits into the new data platform. Key Takeaways: 00:00 Introduction.02:11 Sanket's background and role leading data platform initiatives at Trading Technologies.04:32 The data landscape at TT: silos, OLTP-driven reporting, and an acquired company running orchestration on a 17-18 year old ASP.NET stack.10:53 Choosing Airflow and Snowflake as the foundation of the new data platform.11:36 Why Airflow won over legacy schedulers like Autosys plus Informatica, especially past the 2,000-3,000 job mark.16:34 Collapsing thousands of legacy jobs into 15-20 Airflow DAGs.17:38 Argo vs Airflow: community support and ecosystem as the deciding factors.18:20 Why Astronomer over self-managing Airflow on MWAA or Cloud Composer: CI/CD, observability, SLAs, PagerDuty, and billing alerts.24:07 The operational math: why running Airflow globally needs dedicated specialists, and why a generalist cannot handle upgrades.27:00 Migration strategy and the human side: addressing the "why" so teams move off legacy tools willingly.30:31 Where AI fits in: letting customers dialogue with their own data, and using cloud skills with co-pilots to generate boilerplate DAGs.36:00 The honest take on AI coding tools: accelerators, not replacements for engineering judgment. Resources Mentioned: [Trading Technologies](tradingtechnologies.com)[Apache Airflow](airflow.apache.org)[Astronomer](astronomer.io)[Snowflake](snowflake.com) Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #ApacheAirflow

    Consolidating legacy job schedulers with Airflow at Trading Technologies
  8. Jul 23

    Orchestrating AI video intelligence and evaluation pipelines at Firework

    How do you orchestrate AI workflows when LLM outputs are non-deterministic and evaluation costs can quietly exceed compute costs? Shawn Feng, Head of Data at [Firework](firework.com), joins the show to walk through how his team uses Airflow to coordinate ingestion, embeddings, and evaluation pipelines for AI features in a video commerce platform. The conversation covers the technical stack, the challenges of validating LLM output, cost guardrails, and what Shawn wants to see next from the Airflow project. Key Takeaways: 00:00 Introduction.00:48 What Firework does: a video commerce platform bringing short-form shoppable video and livestream experiences directly onto brand websites and apps.02:13 Shawn's team owns the full data and AI platform stack at Firework, from ingestion through BI and applied AI.02:59 Airflow has been the main orchestration layer for batch transformations since the early open source days, and now powers AI workflows like conversational insights, knowledge-base building, and automated eval pipelines.06:45 The AI workflow stack: ingestion of user interaction and content data, transformation into personalized profiles and embeddings, separate eval pipelines dispatched from third-party systems, all on Snowflake with Cortex AI and coordinated by Airflow.08:55 Why Airflow stuck: flexibility, reliability, and clear visibility into dependencies across SQL, Python, and AI tasks.11:27 The deterministic pipeline problem. Airflow assumes predictable input and output. LLM workflows break that assumption, which forces evaluation layers, traceability, regression testing, and feedback loops into the pipeline itself.13:07 Evaluation cost can exceed compute cost. Validating a single output can mean running four or five parallel eval jobs across prompts and configurations, so caching, batching, and selective evaluation become essential.15:58 Building safeguards: alerting and aborting jobs when cost exceeds thresholds, so the platform stays operationally stable while other improvements catch up.18:30 Wishlist: stronger first-class support for AI workflow patterns (evaluation tracking, prompt experimentation, model observability, event-driven AI orchestration) and continued UI/UX improvement in Airflow 3. Resources Mentioned: [Firework](firework.com)[Apache Airflow](airflow.apache.org)[Snowflake Cortex AI](snowflake.com/en/data-cloud/cortex)[Airflow common.ai provider](airflow.apache.org/docs/apache-airflow-providers-common-ai/stable/index.html)Shawn Feng on LinkedIn (https://www.linkedin.com/in/shawnshifeng/) Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #ApacheAirflow

    Orchestrating AI video intelligence and evaluation pipelines at Firework
5
out of 5
21 Ratings

About

Welcome to The Data Flowcast: Mastering Apache Airflow ® for Data Engineering and AI— the podcast where we keep you up to date with insights and ideas propelling the Airflow community forward. Join us each week, as we explore the current state, future and potential of Airflow with leading thinkers in the community, and discover how best to leverage this workflow management system to meet the ever-evolving needs of data engineering and AI ecosystems. Podcast Webpage: https://www.astronomer.io/podcast/

You Might Also Like