Title Information
Title
Efficient Data Analytics Using Speculative Compilation Techniques
Type of Resource (primo)
dissertations
Name: Personal
Name Part
Spiegelberg, Leonhard Franz
Role
Role Term: Text
creator
Name: Personal
Name Part
Schwarzkopf, Malte
Role
Role Term: Text
Advisor
Name: Personal
Name Part
Kraska, Tim
Role
Role Term: Text
Advisor
Name: Personal
Name Part
Krishnamurthi, Shriram
Role
Role Term: Text
Reader
Name: Corporate
Name Part
Brown University. Department of Computer Science
Role
Role Term: Text
sponsor
Origin Information
Copyright Date
2023
Physical Description
Extent
, None p.
digitalOrigin
born digital
Note: thesis
Thesis (Ph. D.)--Brown University, 2023
Genre (aat)
theses
Abstract
Over the last decade, Python emerged as the language of choice for data processing and model building. Yet, the developer productivity of Python comes at the cost of slow execution speed and high demand for resources. Compiling Python to efficient, optimized machine code could help, but writing a general-purpose Python compiler is inherently difficult. Python’s dynamic typing, dynamic dispatch, and object format all impede standard compiler optimizations, and efforts to write compilers over two decades have failed to produce substantial speedups for practical data science workloads. This dissertation introduces two specialized compilers that leverage the high-level structure of a query to compile Python to efficient code for parallel data science workloads. The first exploits the domain-specific program structure of data science pipelines and generates code specialized to a sample of the input data, which allows making assumptions about types and common-case behavior. This results in a novel data-driven compilation approach that produces highly efficient machine code. Our prototype data analytics system, Tuplex, demonstrates that data-driven compilation beats state-of-the-art systems by a factor of 5× to 91× on realistic workloads. We then explore the idea of hyperspecialization, which generates custom code for disjoint data partitions to avoid sampling errors and improve overall efficiency for heterogeneous datasets. In a distributed setting using serverless functions to avoid sampling errors and improve overall efficiency. Building on Tuplex, we show that this next-generation system, Viton, benefits from hyperspecialization for realistic workloads and reduces cost and runtime by a factor of 2× to 3×. Viton uses serverless Lambda functions to achieve massive parallelism at low latencies together with a more advanced data-driven compilation approach backed by hyperspecialization.
Subject (fast) (authorityURI="http://id.worldcat.org/fast", valueURI="http://id.worldcat.org/fast/01892965")
Topic
Big data
Subject
Topic
Query Processing
Subject
Topic
data science
Subject
Topic
Data analytics
Subject (fast) (authorityURI="http://id.worldcat.org/fast", valueURI="http://id.worldcat.org/fast/00871538")
Topic
Compilers (Computer programs)
Language
Language Term (ISO639-2B)
English
Record Information
Record Content Source (marcorg)
RPB
Record Creation Date (encoding="iso8601")
20230602