activity
20162020
most citedSynthetic Data for Social Good

51 citations · 51 across the 1 of their papers we have counts for

collaborators
Showing cs.DBShow all

7 papers · 1 filter

cs.DB2020

SPORES: Sum-Product Optimization via Relational Equality Saturation for Large Scale Linear Algebra

Yisu Remy Wang, Shana Hutchison, Jonathan Leang +2

Machine learning algorithms are commonly specified in linear algebra (LA). LA expressions can be rewritten into more efficient forms, by taking advantage of input properties such a…

cs.DB2019

Capuchin: Causal Database Repair for Algorithmic Fairness

Babak Salimi, Luke Rodriguez, Bill Howe +1

Fairness is increasingly recognized as a critical component of machine learning systems. However, it is the underlying data on which these systems are trained that often reflect di…

cs.DB2018

Database-Agnostic Workload Management

Shrainik Jain, Jiaqi Yan, Thierry Cruane +1

We present a system to support generalized SQL workload analysis and management for multi-tenant and multi-database platforms. Workload analysis applications are becoming more soph…

cs.DB2018

Privacy-Preserving Synthetic Datasets Over Weakly Constrained Domains

Luke Rodriguez, Bill Howe

Techniques to deliver privacy-preserving synthetic datasets take a sensitive dataset as input and produce a similar dataset as output while maintaining differential privacy. These…

cs.DB2018

MobilityMirror: Bias-Adjusted Transportation Datasets

Luke Rodriguez, Babak Salimi, Haoyue Ping +2

We describe customized synthetic datasets for publishing mobility data. Private companies are providing new transportation modalities, and their data is of high value for integrati…

cs.DB2018

Query2Vec: An Evaluation of NLP Techniques for Generalized Workload Analytics

Shrainik Jain, Bill Howe, Jiaqi Yan +1

We consider methods for learning vector representations of SQL queries to support generalized workload analytics tasks, including workload summarization for index selection and pre…