PrestoDB Blog - PrestoDB

Elevating Presto Query Optimization: Leveraging State-of-the-Art Techniques for Improved Performance

By David Simmen, Anant Aneja, Vivek Bharathan, Zachary Blanco, Aditi Pandit & Ethan Zhang March 21, 2024March 21, 2024

Presto, a prominent open-source distributed SQL query engine, has been at the leading edge of high-performance data analytics for over a decade. In analytical data processing, the effectiveness of query optimization is paramount. Over the last half-century, optimizing SQL queries has been a hotbed of research and development, resulting in groundbreaking innovations. This blog post…

2022 PrestoDB Community in Review

By Ali LeClerc December 30, 2022September 14, 2023

Hello Presto enthusiasts! We here at the Presto Outreach Committee are absolutely thrilled to be entering the new year of 2023. It’s hard to believe that another year has passed, but as we reflect on the past year, we can’t help but feel grateful for the amazing growth and progress we’ve seen in the Presto…

Our Presto Credo for the Truly Open Source SQL Query Engine

By Steven Mih, Girish Baliga & Tim Meehan December 8, 2022September 21, 2023

We believe that data analytics should be democratized—and is why we innovate Presto with state-of-the-art database technology. Trusted governance is important to us—and is why we model our project governance and by laws after the Linux Foundation. TO OUR FELLOW DATA ENGINEERS, SOFTWARE DEVELOPERS, AND DATA PLATFORM ENTHUSIASTS: As the use of data analytics and…

Is PrestoDB the most popular Open Source Data Analytics project?

By Ali LeClerc November 30, 2022September 21, 2023

The Presto Foundation is thrilled to announce that today Presto has been awarded “2022 Editors Choice for Top 3 Data and AI Open Source Projects to Watch” from BigDATAwire. Past winners are a true who’s who in the data world including Apache Spark (2020), Apache Kafka (2018), MongoDB (2019), Apache Cassandra, ElasticSearch and Redis (2021)….

Avoid Data Silos in Presto in Meta: the journey from Raptor to RaptorX

By Rongrong Zhong, James Sun & Ke Wang January 28, 2022September 21, 2023

Raptor is a Presto connector (presto-raptor) that is used to power some critical interactive query workloads in Meta (previously Facebook). Though referred to in the ICDE 2019 paper Presto: SQL on Everything, it remains somewhat mysterious to many Presto users because there is no available documentation for this feature. This article will shed some light…

Native Parquet Writer for Presto

By Lu Niu & Zhenxiao Luo June 29, 2021September 21, 2023

Overview With the wide deployment of Presto in a growing number of companies, Presto is used not only for queries, but also for data ingestion and ETL jobs. There is a need to improve Presto’s file writer performance, especially for popular columnar file formats, e.g. Parquet, and ORC. In this article, we introduce the brand…

Presto Foundation and PrestoDB: Our Commitment to the Presto Open Source Community

By Girish Baliga, Tim Meehan, Dipti Borkar, Zhenxiao Luo, Steven Mih & Bin Fan June 14, 2021September 21, 2023

We recently wrapped up an amazing PrestoCon Day attended by over 600 people from across the globe. The technical discussions and the panel was a clear indication of the growing community. We showcased a number of features contributed by various companies that continue to advance the mission of Presto open source, reiterating our commitment to…

RaptorX: Building a 10X Faster Presto

By James Sun, Ke Wang, Rohit Jain, Saksham Sachdev, Shixuan Fan, Bin Fan, Zhenxiao Luo & Lu Niu February 4, 2021September 21, 2023

RaptorX is an internal project name aiming to boost query latency significantly beyond what vanilla Presto is capable of. This blog post introduces the hierarchical cache work, which is the key building block for RaptorX. With the support of the cache, we are able to boost query performance by 10X. This new architecture can beat…

2020 Recap – A Year with Presto

By Dipti Borkar January 12, 2021September 21, 2023

Tl;dr: 2020 was a huge year for the Presto community. We held our first major conference, PrestoCon, the biggest Presto event ever. We had a massive expansion of our meetup groups with more than 20 sessions held throughout the year, and significant innovations were contributed to Presto! This year has certainly been unique, to say…

Table Scan: Doing The Right Thing With Structured Types

By Orri Erling September 26, 2019September 21, 2023

In the previous article we saw what gains are possible when filtering early and in the right order. In this article we look at how we do this with nested and structured types. We use the 100G TPC-H dataset, but now we group top level columns into structs or maps. Maps, lists and structs are…

Complete Table Scan: A Quantitative Assessment

By Orri Erling July 29, 2019September 21, 2023

In the previous article we looked at the abstract problem statement and possibilities inherent in scanning tables. In this piece we look at the quantitative upside with Presto. We look at a number of queries and explain the findings. The initial impulse motivating this work is the observation that table scan is by far the…

Everything You Always Wanted To Do in Table Scan

By Orri Erling, Maria Basmanova, Ying Su, Tim Meehan & Elon Azoulay June 29, 2019September 21, 2023

Table scan, on the face of it, sounds trivial and boring. What’s there in just reading a long bunch of records from first to last? Aren’t indexing and other kinds of physical design more interesting? As data has gotten bigger, the columnar table scan has only gotten more prominent. The columnar scan is a fairly…

Introducing the Presto blog

By Orri Erling June 28, 2019September 21, 2023

Presto is a key piece of data infrastructure at many companies. The community has many ongoing projects for taking it to new levels of performance and functionality plus unique experience and insight into challenges of scale. We are opening this blog as an informal channel for discussing our work as well as technology trends and…