From Ingestion to Insights:
Building a Data Lakehouse
Pipeline for GitHub Analytics
A 90-minute live masterclass in modern data engineering and big data analytics. Ingest real-world GitHub Archive big data with Apache Airflow & Apache Spark, build ACID lakehouse tables with Apache Hudi, run sub-second queries on the Presto open source distributed SQL query engine, and visualize live metrics in Apache Superset.
Reserve Your Free Spot
Workshop Curriculum: What You'll Learn & Build
Master the complete open lakehouse stack through 7 hands-on lab modules, from S3 storage foundation and Airflow orchestration to Apache Hudi ACID table engineering, sub-second PrestoDB queries, and live Superset dashboards.
S3 Lakehouse Storage Foundation
MinIO S3 • S3A Buckets & HMAC KeysDeploy local S3-compatible cloud object storage, configure S3A endpoint policies, and secure credentials without vendor lock-in.
github-raw-data bucket via MinIO Client (mc)Ingestion & DAG Orchestration
Airflow 2.7 • GitHub Archive Batch DAGSchedule automated batch workflows to pull, validate, and stage 24 hours of GitHub Archive JSON.gz dumps directly into MinIO.
github_events_ingestion DAG with retry & Slack alertsACID Tables with Spark & Hudi
Spark 3.5 + Hudi 0.15 • HoodieStreamerEngineer high-throughput transactional lakehouse tables with Avro schema enforcement, primary keys, and deduplication on object storage.
HoodieStreamer with github-schema.avsc & type partitioningHudi ACID Superpowers & Time-Travel
Python Script • Upserts, GDPR & TimelineMaster mutable big data operations: in-place state updates without table rewrites, GDPR point deletes, and timeline-based time-travel.
hudi_acid_superpowers_demo.py to mutate PR recordsCentralized Catalog with HMS
Hive Metastore 3.1 • Thrift :9083 & MySQL 8Decouple storage from compute engines with a unified Thrift catalog and automated dynamic partition discovery via Hive sync.
Sub-Second Distributed SQL
PrestoDB 0.299 • 3-Node MPP ClusterExecute interactive distributed SQL directly on partitioned Hudi tables with vectorized columnar execution and 5 quality gates.
catalog/hudi.properties & run 5 PyHive gate testsLive Lakehouse BI & Interactive Dashboards
Apache Superset 6.1 • PyHive Driver & Interactive MetricsConnect open-source BI to PrestoDB via PyHive, build virtual SQL views for developer velocity, and publish interactive dashboards with instant slice-and-dice.
End-to-End Lakehouse Flow: From Ingestion to Sub-Second SQL
Experience the production big data pipeline: stage raw GitHub Archive hourly dumps via Apache Airflow, ingest and commit into Apache Hudi ACID tables using Apache Spark on MinIO object storage, and execute distributed queries with sub-second latency on the Presto open source query engine.
GitHub Archive
s3a://github-raw-data/github-raw/Spark 3.5 Engine
github-schema.avscHudi on MinIO
github_events (type=...)PrestoDB 0.299
1 Coord + 2 Workers (:8080)Apache Superset
Frequently Asked Questions
Everything you need to know about joining live, workstation prerequisites, session recordings, and lab access.
Will the workshop recording and lab code be available afterward?
Is this workshop completely free of charge?
Do I need a cloud account for MinIO to participate?
What are the system hardware and software prerequisites?
Everything runs locally on your workstation via Docker Compose. Ensure your machine meets these specifications:
- Hardware: 8GB RAM minimum (16GB recommended to run all cluster services concurrently), 15GB free disk space, and 4+ CPU cores (Apple Silicon, Intel x86_64, or AMD64).
- Software: Docker Desktop 4.20+ with Docker Compose v2, Git, curl, and a bash or zsh terminal.
- Foundational Knowledge: Basic SQL querying (no prior experience with Presto, Apache Hudi, or Apache Spark is required).
To verify readiness in your terminal before the workshop, run: $ docker info && git --version
Where can I ask questions if I get stuck during the hands-on lab?
Join the Presto Foundation Community
Connect with thousands of data engineers, core committers, and lakehouse architects across our official community channels.
Stay updated with official Presto Foundation announcements, upcoming events, engineering blogs, and community meetups.
Follow Presto FoundationSlack Community
Directly collaborate with Presto core committers, ask troubleshooting questions, and get peer assistance during and after the workshop.
Join Presto SlackYouTube Channel
Watch recorded PrestoCon sessions, deep-dive technical webinars, architecture masterclasses, and feature walkthroughs on-demand.
Watch on YouTube