activity
20182021
most citedPIDForest: Anomaly Detection via Partial Identification

4 citations · 4 across the 3 of their papers we have counts for

collaborators

6 papers

cs.LG2021

One Network Fits All? Modular versus Monolithic Task Formulations in Neural Networks

Atish Agarwala, Abhimanyu Das, Brendan Juba +4

Can deep learning solve multiple tasks simultaneously, even when they are unrelated and very different? We investigate how the representations of the underlying tasks affect the ab…

cs.LG2021

Multicalibrated Partitions for Importance Weights

Parikshit Gopalan, Omer Reingold, Vatsal Sharan +1

The ratio between the probability that two distributions and give to points are known as importance weights or propensity scores and play a fundamental role in many dif…

cs.LG20194 cited

PIDForest: Anomaly Detection via Partial Identification

Parikshit Gopalan, Vatsal Sharan, Udi Wieder

We consider the problem of detecting anomalies in a large dataset. We propose a framework called Partial Identification which captures the intuition that anomalies are easy to dist…

cs.LG2019

Memory-Sample Tradeoffs for Linear Regression with Small Error

Vatsal Sharan, Aaron Sidford, Gregory Valiant

We consider the problem of performing linear regression over a stream of -dimensional examples, and show that any algorithm that uses a subquadratic amount of memory exhibits a…

cs.LG2018

Efficient Anomaly Detection via Matrix Sketching

Vatsal Sharan, Parikshit Gopalan, Udi Wieder

We consider the problem of finding anomalies in high-dimensional data using popular PCA based anomaly scores. The naive algorithms for computing these scores explicitly compute the…

cs.DB2018

Moment-Based Quantile Sketches for Efficient High Cardinality Aggregation Queries

Edward Gan, Jialin Ding, Kai Sheng Tai +2

Interactive analytics increasingly involves querying for quantiles over sub-populations of high cardinality datasets. Data processing engines such as Druid and Spark use mergeable…