LIVE HANDS-ON WORKSHOP
    October 15, 2026 • 8:30 AM PDT / 11:30 AM EDT / 9:00 PM IST

    From Ingestion to Insights:
    Building a Data Lakehouse
    Pipeline for GitHub Analytics

    A 90-minute live masterclass in modern data engineering and big data analytics. Ingest real-world GitHub Archive big data with Apache Airflow & Apache Spark, build ACID lakehouse tables with Apache Hudi, run sub-second queries on the Presto open source distributed SQL query engine, and visualize live metrics in Apache Superset.

    Yi-Hong Wang - Senior Software Developer at IBM
    WORKSHOP INSTRUCTOR
    Yi-Hong Wang
    Senior Software Developer @IBM

    Reserve Your Free Spot

    10 DAYS
    :
    03 HOURS
    :
    25 MINS
    :
    48 SECS
    TECHNOLOGIES YOU'LL MASTER IN THIS WORKSHOP
    Apache Airflow Data Engineering Workflow Orchestration
    Apache Spark Big Data Processing Engine
    Apache Hudi Open Source Lakehouse ACID Table Format
    MinIO S3-Compatible Object Storage
    Apache Hive Metastore Centralized Catalog
    Presto Open Source Distributed SQL Query Engine
    Apache Superset Open Source Data Analytics & BI

    Workshop Curriculum: What You'll Learn & Build

    Master the complete open lakehouse stack through 7 hands-on lab modules, from S3 storage foundation and Airflow orchestration to Apache Hudi ACID table engineering, sub-second PrestoDB queries, and live Superset dashboards.

    MODULE 01 Storage
    MinIO Object Storage

    S3 Lakehouse Storage Foundation

    MinIO S3 • S3A Buckets & HMAC Keys

    Deploy local S3-compatible cloud object storage, configure S3A endpoint policies, and secure credentials without vendor lock-in.

    Lab: Provision github-raw-data bucket via MinIO Client (mc)
    Goal: Verified S3 object storage bucket ready for raw archive staging
    MODULE 02 Orchestration
    Apache Airflow Orchestration

    Ingestion & DAG Orchestration

    Airflow 2.7 • GitHub Archive Batch DAG

    Schedule automated batch workflows to pull, validate, and stage 24 hours of GitHub Archive JSON.gz dumps directly into MinIO.

    Lab: Deploy github_events_ingestion DAG with retry & Slack alerts
    Goal: Automated batch ingestion pipeline staging 24h raw archive dumps (~3.6M events)
    MODULE 03 Table Format
    Apache Spark Engine

    ACID Tables with Spark & Hudi

    Spark 3.5 + Hudi 0.15 • HoodieStreamer

    Engineer high-throughput transactional lakehouse tables with Avro schema enforcement, primary keys, and deduplication on object storage.

    Lab: Run HoodieStreamer with github-schema.avsc & type partitioning
    Goal: Optimized Copy-on-Write Hudi lakehouse table on MinIO
    MODULE 04 ACID Superpowers
    Apache Hudi ACID Superpowers

    Hudi ACID Superpowers & Time-Travel

    Python Script • Upserts, GDPR & Timeline

    Master mutable big data operations: in-place state updates without table rewrites, GDPR point deletes, and timeline-based time-travel.

    Lab: Run hudi_acid_superpowers_demo.py to mutate PR records
    Goal: Verified in-place record upserts and compliant point deletes
    MODULE 05 Catalog
    Apache Hive Metastore

    Centralized Catalog with HMS

    Hive Metastore 3.1 • Thrift :9083 & MySQL 8

    Decouple storage from compute engines with a unified Thrift catalog and automated dynamic partition discovery via Hive sync.

    Lab: Configure HMS Thrift service on port 9083 backed by MySQL 8
    Goal: Centralized schema & partition metadata catalog
    MODULE 06 Distributed SQL
    PrestoDB Distributed Query Engine

    Sub-Second Distributed SQL

    PrestoDB 0.299 • 3-Node MPP Cluster

    Execute interactive distributed SQL directly on partitioned Hudi tables with vectorized columnar execution and 5 quality gates.

    Lab: Configure catalog/hudi.properties & run 5 PyHive gate tests
    Goal: Sub-second query engine with zero disk spill across 3 nodes

    End-to-End Lakehouse Flow: From Ingestion to Sub-Second SQL

    Experience the production big data pipeline: stage raw GitHub Archive hourly dumps via Apache Airflow, ingest and commit into Apache Hudi ACID tables using Apache Spark on MinIO object storage, and execute distributed queries with sub-second latency on the Presto open source query engine.

    1. Ingest
    ➔
    2. Batch
    ➔
    3. Lakehouse
    ➔
    4. Presto SQL
    ➔
    5. Superset
    1. INGESTION

    GitHub Archive

    PushEvent • 4,120 rows
    Staged to MinIO: s3a://github-raw-data/github-raw/
    Apache Spark
    2. PROCESSING

    Spark 3.5 Engine

    HoodieStreamer 10K ev/sec
    Avro Schema: github-schema.avsc
    Apache Hudi
    3. LAKEHOUSE

    Hudi on MinIO

    Commit: 20261015090100
    Table & Partition: github_events (type=...)
    PrestoDB
    4. DISTRIBUTED SQL

    PrestoDB 0.299

    3 Nodes Online
    Cluster Topology: 1 Coord + 2 Workers (:8080)
    Apache Superset
    5. BI & DASHBOARDS

    Apache Superset

    Live Dashboards Updated
    3.6M+ Events Processed
    PHASE 1/5 AIRFLOW
    Polling data.gharchive.org • Staging hourly JSON.gz to s3a://github-raw-data/github-raw/ ...
    Elapsed: 0.00s
    INGESTION VELOCITY
    3.6M events/batch
    24-Hour GitHub Archive batch ingestion
    MINIO READ THROUGHPUT
    2.91 GB/s
    Bloom filter file slice pruning
    CLUSTER TOPOLOGY
    3 / 3 Nodes Active
    0 bytes disk spill (In-Memory SIMD)
    QUERY EXECUTION TIME
    380 ms
    Sub-second vectorized aggregations

    Frequently Asked Questions

    Everything you need to know about joining live, workstation prerequisites, session recordings, and lab access.

    Will the workshop recording and lab code be available afterward?
    Yes! All registered attendees will receive the full recording, slides, and access to the complete GitHub repository containing all Docker Compose files, configuration templates, and setup guides within 24 hours of the live broadcast.
    Is this workshop completely free of charge?
    Yes, 100% free. This is an official community education initiative powered by the Presto Foundation and the open-source community to advance lakehouse engineering skills.
    Do I need a cloud account for MinIO to participate?
    No. MinIO is high-performance, open-source, and runs locally via Docker. You can run the entire lakehouse lab completely on your local workstation without needing any paid cloud accounts or incurring any cloud bills.
    What are the system hardware and software prerequisites?

    Everything runs locally on your workstation via Docker Compose. Ensure your machine meets these specifications:

    • Hardware: 8GB RAM minimum (16GB recommended to run all cluster services concurrently), 15GB free disk space, and 4+ CPU cores (Apple Silicon, Intel x86_64, or AMD64).
    • Software: Docker Desktop 4.20+ with Docker Compose v2, Git, curl, and a bash or zsh terminal.
    • Foundational Knowledge: Basic SQL querying (no prior experience with Presto, Apache Hudi, or Apache Spark is required).

    To verify readiness in your terminal before the workshop, run: $ docker info && git --version

    Where can I ask questions if I get stuck during the hands-on lab?
    You will have direct interactive access to Yi-Hong Wang during the dedicated live Q&A, plus continuous peer and maintainer support in the official Presto Community Slack workspace.

    Ready to Master the Modern Open Data Lakehouse?

    Join hundreds of data engineers, big data architects, and analytics leaders in this live, deeply technical masterclass with Presto, Apache Hudi, Spark, and Airflow. Reserve your free spot before registration closes.

    RESERVE YOUR FREE SPOT NOW