7 papers · 1 filter
How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions
Jeff A. Bilmes, Gantavya Bhatt, Arnav M. Das
Neural scaling laws appraise data through dataset size, while the Vendi Score uses quantum entropy to measure dataset value. We show both that common neural-scaling-law objectives…
COBRA: COmBinatorial Retrieval Augmentation for Few-Shot Adaptation
Arnav M. Das, Gantavya Bhatt, Lilly Kumari +2
Retrieval augmentation, the practice of retrieving additional data from large auxiliary pools, has emerged as an effective technique for enhancing model performance in the low-data…
Deep Submodular Peripteral Networks
Gantavya Bhatt, Arnav Das, Jeff Bilmes
Submodular functions, crucial for various applications, often lack practical learning methods for their acquisition. Seemingly unrelated, learning a scaling from oracles offering g…
Effective Backdoor Mitigation in Vision-Language Models Depends on the Pre-training Objective
Sahil Verma, Gantavya Bhatt, Avi Schwarzschild +6
Despite the advanced capabilities of contemporary machine learning (ML) models, they remain vulnerable to adversarial and backdoor attacks. This vulnerability is particularly conce…
LabelBench: A Comprehensive Framework for Benchmarking Adaptive Label-Efficient Learning
Jifan Zhang, Yifang Chen, Gregory Canal +8
Labeled data are critical to modern machine learning applications, but obtaining labels can be expensive. To mitigate this cost, machine learning methods, such as transfer learning…
Accelerating Batch Active Learning Using Continual Learning Techniques
Arnav Das, Gantavya Bhatt, Megh Bhalerao +3
A major problem with Active Learning (AL) is high training costs since models are typically retrained from scratch after every query round. We start by demonstrating that standard…