45 citations · 69 across the 4 of their papers we have counts for
3 papers · 1 filter
DataHub: Collaborative Data Science & Dataset Version Management at Scale
Anant Bhardwaj, Souvik Bhattacherjee, Amit Chavan +4
Relational databases have limited support for data collaboration, where teams collaboratively curate and analyze large datasets. Inspired by software version control systems like g…
MDCC: Multi-Data Center Consistency
Tim Kraska, Gene Pang, Michael J. Franklin +1
Replicating data across multiple data centers not only allows moving the data closer to the user and, thus, reduces latency for applications, but also increases the availability in…
BlinkDB: Queries with Bounded Errors and Bounded Response Times on Very Large Data
Sameer Agarwal, Aurojit Panda, Barzan Mozafari +2
In this paper, we present BlinkDB, a massively parallel, sampling-based approximate query engine for running ad-hoc, interactive SQL queries on large volumes of data. The key insig…