activity
20182020
most citedSherlock: A Deep Learning Approach to Semantic Data Type Detection

12 citations · 15 across the 2 of their papers we have counts for

collaborators

5 papers

cs.LG2020

ARDA: Automatic Relational Data Augmentation for Machine Learning

Nadiia Chepurko, Ryan Marcus, Emanuel Zgraggen +3

Automatic machine learning (\AML) is a family of techniques to automate the process of training predictive models, aiming to both improve performance and make machine learning more…

cs.LG201912 cited

Sherlock: A Deep Learning Approach to Semantic Data Type Detection

Madelon Hulsebos, Kevin Hu, Michiel Bakker +5

Correctly detecting the semantic type of data columns is crucial for data science tasks such as automated data cleaning, schema matching, and data discovery. Existing data preparat…

cs.HC20193 cited

VizNet: Towards A Large-Scale Visualization Learning and Benchmarking Repository

Kevin Hu, Neil Gaikwad, Michiel Bakker +7

Researchers currently rely on ad hoc datasets to train automated visualization tools and evaluate the effectiveness of visualization designs. These exemplars often lack the charact…

cs.CR2019

STAR: Statistical Tests with Auditable Results

Sacha Servan-Schreiber, Olga Ohrimenko, Tim Kraska +1

We present STAR: a novel system aimed at solving the complex issue of "p-hacking" and false discoveries in scientific studies. STAR provides a concrete way for ensuring the applica…

cs.DB2018

IDEBench: A Benchmark for Interactive Data Exploration

Philipp Eichmann, Carsten Binnig, Tim Kraska +1

Existing benchmarks for analytical database systems such as TPC-DS and TPC-H are designed for static reporting scenarios. The main metric of these benchmarks is the performance of…