The Data Flowcast: Mastering Apache Airflow ® for Data Engineering and AI

Astronomer

Welcome to The Data Flowcast: Mastering Apache Airflow ® for Data Engineering and AI— the podcast where we keep you up to date with insights and ideas propelling the Airflow community forward. Join us each week, as we explore the current state, future and potential of Airflow with leading thinkers in the community, and discover how best to leverage this workflow management system to meet the ever-evolving needs of data engineering and AI ecosystems. Podcast Webpage: https://www.astronomer.io/podcast/

  1. 1d ago

    How Airflow orchestration decisions impact Spark performance

    Airflow and Spark are one of the most common combinations in modern data platforms, but the orchestration decisions made on the Airflow side often determine whether Spark jobs run fast and cheap or slow and expensive. In this episode, Kenten is joined by [Meni Shmueli](linkedin.com/in/meni-shmueli-dataflint), Co-Founder and CEO at [DataFlint](dataflint.io), to dig into how Airflow and Spark fit together, where teams go wrong, and how AI is changing the way they reason about cost and performance. Key Takeaways: 00:00 Introduction.01:40 Meni's background as a data engineer and what led him to start DataFlint, a production observability agent for Apache Spark.02:53 Why the Airflow plus Spark combination is so common, and how each tool plays to its strengths.04:23 How orchestration decisions in Airflow directly impact Spark performance and cost.04:44 A customer story where parallelizing Airflow tasks made Spark jobs slower, less stable, and more expensive, and the opposite case where sequential runs left compute on the table.07:02 The number one mistake teams make benchmarking pipelines: only looking at the Airflow side and ignoring underlying Spark cost and resource usage.08:13 What DataFlint's Airflow and Astro integration gives teams, and how it brings production context into AI agents.10:46 How AI is changing pipeline optimization, including holistic scheduling across hundreds of pipelines and connecting context from Spark, FinOps, and cloud.12:21 A customer migration from Databricks to EMR that cut workflow costs by 80%, with examples of up to 100x optimizations.14:23 Where the Airflow and Spark story could be better, including Spark Declarative Pipelines and tighter feedback between the two projects. Resources Mentioned: [Orchestrate Everything](https://astronomer.link/data-flowcast-oe)[DataFlint](dataflint.io)[Apache Airflow](airflow.apache.org)[Apache Spark](spark.apache.org)[Astro](astronomer.io/astro)[Airflow Summit](airflowsummit.org) Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #ApacheAirflow

    How Airflow orchestration decisions impact Spark performance
  2. Aug 6

    Using Airflow for diverse client projects at Accion Labs

    When a real-time loan eligibility scoring pipeline is built on cron jobs, midnight pages are inevitable. In this episode, [Chandan Gowda](linkedin.com/in/chandan-gowda-a-h-744908195), Data Engineer at [Accion Labs](accionlabs.com), joins Kenten to discuss how his team uses Airflow across client projects, including a financial services scoring use case and a POC applying production-grade orchestration to RAG and GenAI data pipelines. Key Takeaways: 00:00 Introduction.01:00 What Accion Labs does as a technology consulting and services firm working across BFSI, healthcare, and retail.02:00 Chandan's role at the intersection of data engineering and GenAI, building pipelines one week and RAG-based agents the next.04:20 Why Airflow tends to win client evaluations: infrastructure agnostic, no cloud lock-in, fine-grained control over pipeline logic.06:00 Containerizing Airflow on Kubernetes or VMs so migrations between clouds don't require a rewrite.08:14 The loan eligibility scoring use case for a financial services client, and replacing fragile cron jobs with a single Airflow DAG end to end, cutting effort by about 25%.10:42 Triggering strategy: S3 file sensors as the primary trigger handling 90% of runs, plus a scheduled fallback as a safety net.12:53 End-to-end flow inside Airflow: ingestion into the data lake, validation and transformation with credit bureau joins, containerized model inference, and writeback to the loan management system.15:56 The AI orchestration POC and why the data feeding GenAI models needs the same rigor as any production pipeline.18:09 Using Airflow to detect document changes and re-chunk and re-embed only what changed, with quality thresholds and rollback before promoting to live.20:23 The roadmap: model evaluation pipelines, multi-agent orchestration, and provider packages for LangChain, OpenAI, and Hugging Face.22:32 Wishlist for Airflow: native event-driven triggers beyond polling sensors, first-class observability for AI workloads, and better dynamic DAG generation at scale. Resources Mentioned: [Orchestrate Everything](https://astronomer.link/data-flowcast-oe)[Accion Labs](accionlabs.com)[Apache Airflow](airflow.apache.org)[Airflow LangChain provider](airflow.apache.org/docs/apache-airflow-providers-langchain/stable/index.html)[Airflow OpenAI provider](airflow.apache.org/docs/apache-airflow-providers-openai/stable/index.html) Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #Airflow

    Using Airflow for diverse client projects at Accion Labs
  3. Jul 30

    Consolidating legacy job schedulers with Airflow at Trading Technologies

    Migrating two decades of legacy job orchestration is the kind of project most teams quietly avoid. In this episode, Marc Lamberti is joined by [Sanket Patel](linkedin.com/in/sanket-patel-a2a57b3a), Director of Engineering at Trading Technologies, to talk through how his team is consolidating 20 years of C#/.NET-based scheduling and a recently acquired company's stack onto Airflow, Astronomer, and Snowflake. The conversation covers orchestrator selection, observability, the operational economics of managed Airflow, the human side of migrations, and where AI fits into the new data platform. Key Takeaways: 00:00 Introduction.02:11 Sanket's background and role leading data platform initiatives at Trading Technologies.04:32 The data landscape at TT: silos, OLTP-driven reporting, and an acquired company running orchestration on a 17-18 year old ASP.NET stack.10:53 Choosing Airflow and Snowflake as the foundation of the new data platform.11:36 Why Airflow won over legacy schedulers like Autosys plus Informatica, especially past the 2,000-3,000 job mark.16:34 Collapsing thousands of legacy jobs into 15-20 Airflow DAGs.17:38 Argo vs Airflow: community support and ecosystem as the deciding factors.18:20 Why Astronomer over self-managing Airflow on MWAA or Cloud Composer: CI/CD, observability, SLAs, PagerDuty, and billing alerts.24:07 The operational math: why running Airflow globally needs dedicated specialists, and why a generalist cannot handle upgrades.27:00 Migration strategy and the human side: addressing the "why" so teams move off legacy tools willingly.30:31 Where AI fits in: letting customers dialogue with their own data, and using cloud skills with co-pilots to generate boilerplate DAGs.36:00 The honest take on AI coding tools: accelerators, not replacements for engineering judgment. Resources Mentioned: [Trading Technologies](tradingtechnologies.com)[Apache Airflow](airflow.apache.org)[Astronomer](astronomer.io)[Snowflake](snowflake.com) Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #ApacheAirflow

    Consolidating legacy job schedulers with Airflow at Trading Technologies
  4. Jul 23

    Orchestrating AI video intelligence and evaluation pipelines at Firework

    How do you orchestrate AI workflows when LLM outputs are non-deterministic and evaluation costs can quietly exceed compute costs? Shawn Feng, Head of Data at [Firework](firework.com), joins the show to walk through how his team uses Airflow to coordinate ingestion, embeddings, and evaluation pipelines for AI features in a video commerce platform. The conversation covers the technical stack, the challenges of validating LLM output, cost guardrails, and what Shawn wants to see next from the Airflow project. Key Takeaways: 00:00 Introduction.00:48 What Firework does: a video commerce platform bringing short-form shoppable video and livestream experiences directly onto brand websites and apps.02:13 Shawn's team owns the full data and AI platform stack at Firework, from ingestion through BI and applied AI.02:59 Airflow has been the main orchestration layer for batch transformations since the early open source days, and now powers AI workflows like conversational insights, knowledge-base building, and automated eval pipelines.06:45 The AI workflow stack: ingestion of user interaction and content data, transformation into personalized profiles and embeddings, separate eval pipelines dispatched from third-party systems, all on Snowflake with Cortex AI and coordinated by Airflow.08:55 Why Airflow stuck: flexibility, reliability, and clear visibility into dependencies across SQL, Python, and AI tasks.11:27 The deterministic pipeline problem. Airflow assumes predictable input and output. LLM workflows break that assumption, which forces evaluation layers, traceability, regression testing, and feedback loops into the pipeline itself.13:07 Evaluation cost can exceed compute cost. Validating a single output can mean running four or five parallel eval jobs across prompts and configurations, so caching, batching, and selective evaluation become essential.15:58 Building safeguards: alerting and aborting jobs when cost exceeds thresholds, so the platform stays operationally stable while other improvements catch up.18:30 Wishlist: stronger first-class support for AI workflow patterns (evaluation tracking, prompt experimentation, model observability, event-driven AI orchestration) and continued UI/UX improvement in Airflow 3. Resources Mentioned: [Firework](firework.com)[Apache Airflow](airflow.apache.org)[Snowflake Cortex AI](snowflake.com/en/data-cloud/cortex)[Airflow common.ai provider](airflow.apache.org/docs/apache-airflow-providers-common-ai/stable/index.html)Shawn Feng on LinkedIn (https://www.linkedin.com/in/shawnshifeng/) Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #ApacheAirflow

    Orchestrating AI video intelligence and evaluation pipelines at Firework
  5. Jul 16

    Orchestrating Retail Data Pipelines at Saks Global

    Saks Global runs one of the largest retail data operations in the US, with around 8 million SKUs flowing across point-of-sale, e-commerce, catalog, and fraud detection systems into Snowflake. In this episode, [Shailesh Kadam](linkedin.com), Architect at [Saks Global](saks.com), joins Kenten to walk through how Airflow acts as the nervous system tying it all together, why they moved from self-managed Kubernetes to Astro, what is driving their Airflow 3 upgrade, and how they are approaching agentic AI, MCP, and credential security. Key Takeaways: 00:00 Introduction.01:31 Saks Global today. Shailesh describes the business after separating e-commerce from brick and mortar and acquiring Neiman Marcus, and the modern cloud-native stack on AWS, Snowflake, and Airflow.02:50 8 million SKUs in motion. Why every name, image, inventory, and price change has to flow in near real time across operational systems.04:50 What the pipelines look like. Point-of-sale ingestion, fraud signals to third parties like Fiserv, and hourly product catalog feeds out to Meta and Google.07:30 Moving off self-managed Kubernetes to Astro. Shailesh contrasts past experience with Kubernetes, IBM Tivoli, and Control-M against running on Astro.09:35 Upgrading to Airflow 3. Event and asset-based scheduling, DAG versioning, task isolation, and using Otto to convert DAGs in a phased rollout.13:13 Agentic AI and MCP on the roadmap. How Saks plans to use Airflow's MCP for LLM-driven product classification and to feed Snowflake analyses like churn and spend.18:01 Securing PII and credentials. Secrets backends, cloud secret manager integration, key rotation, and keeping credentials out of DAG code.21:02 Wishlist for Airflow. Interactive data lineage across DAGs and a UI-based debugging interface for support teams. Resources Mentioned: [Apache Airflow](airflow.apache.org)[Astro](astronomer.io/product)[Otto, the Astronomer data engineering agent](astronomer.io)[Snowflake](snowflake.com)[Saks Global](saks.com) Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #Airflow

    Orchestrating Retail Data Pipelines at Saks Global
  6. Jul 9

    What's New in Apache Airflow® 3.3

    Airflow 3.3 is here, with a set of features to help with the messy realities of production pipelines: persisting state across retries, reacting intelligently to different failure types, and partitioning assets by more than just time. In this episode, Marc Lamberti, Education Content Lead at [Astronomer](astronomer.io), joins Kenten Danas to walk through what's new in the release and where each feature actually pays off. Key Takeaways: 00:00 Introduction.01:46 The new task state store (AIP-103) lets tasks persist state across retries, so a long-running Spark job can be reattached after a worker failure instead of being duplicated on retry.03:46 The asset state store enables watermarking patterns: persist the last processed date or offset to an asset and resume from there on the next run.05:33 Why this matters for agentic workflows: resume an agent from where it left off rather than replaying every action.06:58 Why XComs don't solve this problem: they get reinitialized on every retry.09:27 Pluggable retries let you attach a retry policy to a task that branches on the exception type. Retry on transient errors, stop immediately on a 403.11:42 Subclassing the retry rule for more complex logic, including dynamic retry counts that used to require hacking the metadatabase.15:52 Updates to asset partitions in 3.3: segment-based partitioning with fan-out and roll-up mappers for downstream DAGs.21:42 Running tasks in Java and Go, moving Airflow toward a multi-language orchestrator.23:33 DAG versioning improvement: choose whether a manual rerun uses the most recent DAG version or the original version from that run.25:03 Advice for teams still on Airflow 2: use the upgrade ebook and Astro's AI migration tooling to handle the undifferentiated heavy lifting. Resources Mentioned: AstronomerAirflow 3.3 Release NotesThe Task State StoreUpdates to the asset partitions featureRetry policiesMulti-language support3.3 WebinarUpgrading from Airflow 2 to 3 ebook Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #Airflow

    What's New in Apache Airflow® 3.3
  7. Jun 25

    Running Airflow 3 in a regulated environment at OTPP

    Running Apache Airflow at a major pension fund means balancing strict compliance requirements with the need to move fast on new capabilities. On this episode, Kowsy Narayan, Cloud Data Platform Lead, Data Engineering at [Ontario Teachers' Pension Plan](otpp.com), joins host Kenten Danas to walk through OTPP's cloud migration, their move to Airflow 3, and going fully live on remote execution. Key Takeaways: 00:00 Introduction.01:18 Inside the OTPP data platform team and what they're responsible for across cloud migration, standards, and enablement.02:33 What's driving OTPP's multi-year move off on-prem to a cloud architecture built around scalability and resilience.02:57 The new stack: Snowflake as the enterprise data platform, dbt for transformation, and Airflow as the orchestrator in the middle.04:15 Why OTPP chose Astronomer: active contributions to the Airflow OSS project, fast runtime releases, and built-in monitoring, observability, and RBAC.05:50 Evolving from dbt core with Bash operators to dbt Cosmos for model-level granularity, lineage, and precise failure recovery, plus a performance boost from watcher mode.08:00 Upgrading from Airflow 2.9 to Airflow 3, using the Astro CLI and linters to catch deprecations quickly.09:32 The drivers behind adopting remote execution: keeping data inside the security perimeter and scaling workloads on their own Kubernetes cluster.11:35 How remote execution replaced a complex network architecture of VPN tunnels and firewall rules, removing latency along the way.12:53 The POV process, success criteria, and a six week timebox to validate remote execution before going to production.14:14 Going fully live: OTPP's last hosted deployment was sunset just before recording.15:06 What Kowsy wants next from Airflow: AI orchestration capabilities and continued maturation of remote execution. Resources Mentioned: [Ontario Teachers' Pension Plan](otpp.com)[Apache Airflow](airflow.apache.org)[Astronomer](astronomer.io)[Cosmos](astronomer.io/cosmos) Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #Airflow

    Running Airflow 3 in a regulated environment at OTPP
  8. Jun 11

    Managing a Customer Analytics Platform with Airflow at Skimlinks

    Skimlinks runs a reporting platform that serves around 2,000 weekly publisher users, and the data infrastructure behind it runs on Airflow. In this episode, Julian Larralde, Director of Data Engineering at Skimlinks, walks through the stack, the migration from external task sensors to event-driven Assets, and a YAML-based DAG factory the team built to onboard new publishers without rewriting Python.  Key Takeaways 00:00 Introduction.00:45 What Skimlinks does and how it operates as an affiliate marketing network aggregator for publishers.02:12 Julian's team and the data platform they own: a reporting portal that serves ~2,000 weekly publisher users.03:07 The stack: real-time ingestion into BigQuery, Airflow as the orchestrator, raw / silver / gold layers, and Apache Druid as the serving database for sub-second BI queries.04:50 Reusing the same data marts for ~100 internal customers across marketing, finance, operations, and account management.06:25 Airflow as the single orchestrator: BigQuery operators for SQL business logic, plus raw file exports for the largest publishers.08:08 Moving from external task sensors to datasets (now Assets) and what the migration actually solved.09:18 Why sensor polling created scheduler load and worker overload, and how event-driven Assets fixed both.10:15 The lineage view in the Airflow UI that came as a bonus after the Assets migration.10:49 The vision for multi-tenant Airflow inside Skimlinks: replacing cron, Rundeck, and team-local Airflow instances with a shared platform.14:31 Building a custom DAG factory with YAML configuration for onboarding new publishers.17:33 Breaking a single Python class into single-responsibility components for the DataPipe project.19:07 Adding a Pydantic layer so misconfigured YAML fails at DAG parse time instead of run time.20:31 Using AI assistance to guide refactoring decisions and generate tests across the new class structure.22:34 What Julian wants from Airflow next: asset watchers paired with data contracts. Resources Mentioned Skimlinks - skimlinks.comApache Airflow - airflow.apache.orgAstronomer - astronomer.ioGoogle BigQuery - cloud.google.com/bigqueryApache Druid - druid.apache.orgPydantic - docs.pydantic.devLooker - cloud.google.com/looker Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #Airflow

    Managing a Customer Analytics Platform with Airflow at Skimlinks
5
out of 5
21 Ratings

About

Welcome to The Data Flowcast: Mastering Apache Airflow ® for Data Engineering and AI— the podcast where we keep you up to date with insights and ideas propelling the Airflow community forward. Join us each week, as we explore the current state, future and potential of Airflow with leading thinkers in the community, and discover how best to leverage this workflow management system to meet the ever-evolving needs of data engineering and AI ecosystems. Podcast Webpage: https://www.astronomer.io/podcast/

You Might Also Like