DIVA: Dataset Derivative of a Learning Task
arXiv:2111.09785
Abstract
We present a method to compute the derivative of a learning task with respect to a dataset. A learning task is a function from a training set to the validation error, which can be represented by a trained deep neural network (DNN). The "dataset derivative" is a linear operator, computed around the trained model, that informs how perturbations of the weight of each training sample affect the validation error, usually computed on a separate validation dataset. Our method, DIVA (Differentiable Validation) hinges on a closed-form differentiable expression of the leave-one-out cross-validation error around a pre-trained DNN. Such expression constitutes the dataset derivative. DIVA could be used for dataset auto-curation, for example removing samples with faulty annotations, augmenting a dataset with additional relevant samples, or rebalancing. More generally, DIVA can be used to optimize the dataset, along with the parameters of the model, as part of the training process without the need for a separate validation dataset, unlike bi-level optimization methods customary in AutoML. To illustrate the flexibility of DIVA, we report experiments on sample auto-curation tasks such as outlier rejection, dataset extension, and automatic aggregation of multi-modal data.
References in corpus (16)
- Learning Transferable Visual Models From Natural Language Supervision
- Neural Architecture Search with Reinforcement Learning
- Neural Architecture Search: A Survey
- AutoML: A Survey of the State-of-the-Art
- DARTS: Differentiable Architecture Search
- Learning to Reweight Examples for Robust Deep Learning
- AutoAugment: Learning Augmentation Policies from Data
- Estimating Training Data Influence by Tracing Gradient Descent
- Selection via Proxy: Efficient Data Selection for Deep Learning
- Rethinking the Hyperparameters for Fine-tuning
- DADA: Differentiable Automatic Data Augmentation
- Understanding the role of importance weighting for deep learning
- Gradients as Features for Deep Representation Learning
- A linearized framework and a new benchmark for model selection for fine-tuning
- Hard Sample Mining for the Improved Retraining of Automatic Speech Recognition
- Direct Differentiable Augmentation Search