In-Memory Indexed Caching for Distributed Data Processing
Proceedings - IEEE International Parallel and Distributed Processing Symposium, pp. 104–114
Abstract
Powerful abstractions such as dataframes are only as efficient as their underlying runtime system. The de-facto distributed data processing framework, Apache Spark, is poorly suited for the modern cloud-based data-science workloads due to its outdated assumptions: static datasets analyzed using coarse-grained transformations. In this paper, we introduce the Indexed DataFrame, an in-memory cache that supports a dataframe abstraction which incorporates indexing capabilities to support fast lookup and join operations. Moreover, it supports appends with multi-version concurrency control. We implement the Indexed DataFrame as a lightweight, standalone library which can be integrated with minimum effort in existing Spark programs. We analyze the performance of the Indexed DataFrame in cluster and cloud deployments with real-world datasets and benchmarks using both Apache Spark and Databricks Runtime. In our evaluation, we show that the Indexed DataFrame significantly speeds-up query execution when compared to a non-indexed dataframe, incurring modest memory overhead.
Authors 5
-
Alexandru Uta Aachen
Affiliation as printed
LIACS, Leiden University
-
Affiliation as printed
Databricks
-
University of California, Berkeley
Affiliation as printed
UC Berkeley
-
Delft University of Technology
Affiliation as printed
TU Delft
-
Peter A. Boncz Aachen
Leiden University · Centrum Wiskunde & Informatica · University of California, Berkeley · Delft University of Technology
Affiliation as printed
CWI
LIACS, Leiden University
TU Delft
UC Berkeley
Cited by 2 stored of 2
2 results
Cited by patents worldwide 1 (Lens.org)
-
Data processing method and device and computer readable storage mediumCN116894104A 2023-10-17 Pending
References 61
-
W6774326187details pending0citations