papers

Publications (55)

cs.DB2017

Towards a Holistic Integration of Spreadsheets with Databases: A Scalable Storage Engine for Presentational Data Management

Mangesh Bendre, Vipul Venkataraman, Xinyan Zhou +2

Spreadsheet software is the tool of choice for interactive ad-hoc data management, with adoption by billions of users. However, spreadsheets are not scalable, unlike database syste…

cs.DB2018

Helix: Holistic Optimization for Accelerating Iterative Machine Learning

Doris Xin, Stephen Macke, Litian Ma +3

Machine learning workflow development is a process of trial-and-error: developers iterate on workflows by testing out small modifications until the desired accuracy is achieved. Un…

cs.SE2022

Towards Observability for Production Machine Learning Pipelines

Shreya Shankar, Aditya Parameswaran

Software organizations are increasingly incorporating machine learning (ML) into their product offerings, driving a need for new data management tools. Many of these tools facilita…

cs.LG2018

Helix: Accelerating Human-in-the-loop Machine Learning

Doris Xin, Litian Ma, Jialin Liu +3

Data application developers and data scientists spend an inordinate amount of time iterating on machine learning (ML) workflows -- by modifying the data pre-processing, model train…

cs.DB2016

Squish: Near-Optimal Compression for Archival of Relational Datasets

Yihan Gao, Aditya Parameswaran

Relational datasets are being generated at an alarmingly rapid rate across organizations and industries. Compressing these datasets could significantly reduce storage and archival…

cs.DB2016

SLiMFast: Guaranteed Results for Data Fusion and Source Reliability

Manas Joglekar, Theodoros Rekatsinas, Hector Garcia-Molina +2

We focus on data fusion, i.e., the problem of unifying conflicting data from data sources into a single representation by estimating the source accuracies. We propose SLiMFast, a f…