activity
20142023
most citedIterative Thresholding for Demixing Structured Superpositions in High Dimensions

4 citations · 6 across the 6 of their papers we have counts for

collaborators

17 papers

cs.CV2024

SELECT: A Large-Scale Benchmark of Data Curation Strategies for Image Classification

Benjamin Feuer, Jiawei Xu, Niv Cohen +3

Data curation is the problem of how to collect and organize samples into a dataset that supports efficient learning. Despite the centrality of the task, little work has been devote…

physics.flu-dyn2024

FlowBench: A Large Scale Benchmark for Flow Simulation over Complex Geometries

Ronak Tali, Ali Rabeh, Cheng-Hau Yang +10

Simulating fluid flow around arbitrary shapes is key to solving various engineering problems. However, simulating flow physics across complex geometries remains numerically challen…

cs.LG2024

DIMAT: Decentralized Iterative Merging-And-Training for Deep Learning Models

Nastaran Saadati, Minh Pham, Nasla Saleem +5

Recent advances in decentralized deep learning algorithms have demonstrated cutting-edge performance on various tasks with large pre-trained models. However, a pivotal prerequisite…

cs.CV2024

Mitigating the Impact of Attribute Editing on Face Recognition

Sudipta Banerjee, Sai Pranaswi Mullangi, Shruti Wagle +2

Through a large-scale study over diverse face images, we show that facial attribute editing using modern generative AI models can severely degrade automated face recognition system…

cs.LG2023

Scaling TabPFN: Sketching and Feature Selection for Tabular Prior-Data Fitted Networks

Benjamin Feuer, Chinmay Hegde, Niv Cohen

Tabular classification has traditionally relied on supervised algorithms, which estimate the parameters of a prediction model using its training data. Recently, Prior-Data Fitted N…

cs.CV2023

Exploring Dataset-Scale Indicators of Data Quality

Benjamin Feuer, Chinmay Hegde

Modern computer vision foundation models are trained on massive amounts of data, incurring large economic and environmental costs. Recent research has suggested that improving data…