This Week in Open Lakehouse

Lisa Cao and Scott Haines

This Week in Open Lakehouse is a weekly data engineering podcast discussing the latest in open source news, including query engines, data catalogs, open table formats, orchestrators, and streaming technologies. Produced by the Databricks DevRel team. https://openlakehouse.io/

Episodes

  1. 1d ago

    Iceberg 1.12, OIDC Credentials, Spark Connect Gateway, Arrow JSON Schemas, and Release Management

    In this episode, Lisa and Scott explore the latest developments across the open data ecosystem, including Delta Rust, Apache Iceberg, Apache Spark, Rust-native infrastructure, and the challenges of keeping rapidly evolving technologies interoperable. They discuss release velocity, governance, security, schema evolution, standardization, and the maintenance work required to make open-source innovation reliable in production. Delta Rust 1.0, DataFusion integration, column mapping, schema evolution, and reduced JVM dependenciesApache Iceberg 1.12, variant types, variant shredding, V4 foundations, and performance improvementsSpark security proposals, row-level filters, column masking, trusted execution engines, and policy enforcementOIDC credential propagation, short-lived credentials, identity management, and modern authenticationSpark Connect Gateway, proxy-driven architecture, multi-tenancy, routing, rate limits, and observabilityRust adoption, modular infrastructure, smaller binaries, faster startup times, and composable systemsInteroperability across Spark, Flink, DuckDB, Polars, Arrow, Iceberg, and DeltaSchema evolution, column mapping, positional assumptions, nested data, and logical versus physical representationsStandardized schemas, shared expression models, JSON representations, and reducing translation layersRelease candidates, backports, compatibility matrices, security patches, and the role of open-source maintainers

About

This Week in Open Lakehouse is a weekly data engineering podcast discussing the latest in open source news, including query engines, data catalogs, open table formats, orchestrators, and streaming technologies. Produced by the Databricks DevRel team. https://openlakehouse.io/