Author: Maria Basmanova

Software Engineer at Facebook

Even Faster Unnest
By Ying Su, Maria Basmanova & Orri Erling August 20, 2020September 21, 2023
Unnest is a common operation in Facebook’s daily Presto workload. It converts an ARRAY, MAP, or ROW into a flat relation. Its original implementation used deep copy all the time and was very inefficient. In Unnest Operator Performance Enhancement with Dictionary Blocks, the author improved the Unnest operator by up to 10x in CPU and…
Read More Even Faster Unnest
Improving the Presto planner for better push down and data federation
By Yi He, James Sun, Maria Basmanova, Rongrong Zhong, Jiexi Lin, Saksham Sachdev & Akshay Pall December 23, 2019September 21, 2023
Presto defines a connector API that allows Presto to query any data source that has a connector implementation. The existing connector API provides basic predicate pushdown functionality allowing connectors to perform filtering at the underlying data source. However, there are certain limitations with the existing predicate pushdown functionality that limits what connectors can do. The…
Read More Improving the Presto planner for better push down and data federation
5 design choices—and 1 weird trick — to get 2x efficiency gains in Presto repartitioning
By Ying Su, Orri Erling, Tim Meehan, Sahar Massachi, Bhavani Hari & Maria Basmanova December 20, 2019September 21, 2023
We like Presto. We like it a lot — so much we want to make it better in every way. Here’s an example: we just optimized the PartitionedOutputOperator. It’s now 2-3x more CPU efficient, which, when measured against Facebook’s production workload, translates to 6% gains overall. That’s huge. The optimized repartitioning is in use on…
Read More 5 design choices—and 1 weird trick — to get 2x efficiency gains in Presto repartitioning
Everything You Always Wanted To Do in Table Scan
By Orri Erling, Maria Basmanova, Ying Su, Tim Meehan & Elon Azoulay June 29, 2019September 21, 2023
Table scan, on the face of it, sounds trivial and boring. What’s there in just reading a long bunch of records from first to last? Aren’t indexing and other kinds of physical design more interesting? As data has gotten bigger, the columnar table scan has only gotten more prominent. The columnar scan is a fairly…
Read More Everything You Always Wanted To Do in Table Scan