activity
20152022
most citedPrinciples of Dataset Versioning: Exploring the Recreation/Storage Tradeoff

9 citations · 16 across the 5 of their papers we have counts for

collaborators

7 papers

cs.DB2022

TSEXPLAIN: Explaining Aggregated Time Series by Surfacing Evolving Contributors

Yiru Chen, Silu Huang

Aggregated time series are generated effortlessly everywhere, e.g., "total confirmed covid-19 cases since 2019" and "total liquor sales over time." Understanding "how" and "why" th…

cs.LG2020

Frugal Optimization for Cost-related Hyperparameters

Qingyun Wu, Chi Wang, Silu Huang

The increasing demand for democratizing machine learning algorithms calls for hyperparameter optimization (HPO) solutions at low cost. Many machine learning algorithms have hyperpa…

cs.LG2018

ABC: Efficient Selection of Machine Learning Configuration on Large Dataset

Silu Huang, Chi Wang, Bolin Ding +1

A machine learning configuration refers to a combination of preprocessor, learner, and hyperparameters. Given a set of configurations and a large dataset randomly split into traini…

cs.DB20171 cited

OrpheusDB: Bolt-on Versioning for Relational Databases

Silu Huang, Liqi Xu, Jialin Liu +2

Data science teams often collaboratively analyze datasets, generating dataset versions at each stage of iterative exploration and analysis. There is a pressing need for a system th…

cs.DB2016

Finding Multiple New Optimal Locations in a Road Network

Ruifeng Liu, Ada WaiChee Fu, Zitong Chen +2

We study the problem of optimal location querying for location based services in road networks, which aims to find locations for new servers or facilities. The existing optimal sol…

cs.DB20156 cited

Towards a unified query language for provenance and versioning

Amit Chavan, Silu Huang, Amol Deshpande +3

Organizations and teams collect and acquire data from various sources, such as social interactions, financial transactions, sensor data, and genome sequencers. Different teams in a…