Dataset Distillation with Infinitely Wide Convolutional Networks
arXiv:2107.13034
Abstract
The effectiveness of machine learning algorithms arises from being able to extract useful features from large amounts of data. As model and dataset sizes increase, dataset distillation methods that compress large datasets into significantly smaller yet highly performant ones will become valuable in terms of training efficiency and useful feature extraction. To that end, we apply a novel distributed kernel based meta-learning framework to achieve state-of-the-art results for dataset distillation using infinitely wide convolutional neural networks. For instance, using only 10 datapoints (0.02% of original dataset), we obtain over 65% test accuracy on CIFAR-10 image classification task, a dramatic improvement over the previous best test accuracy of 40%. Our state-of-the-art results extend across many other settings for MNIST, Fashion-MNIST, CIFAR-10, CIFAR-100, and SVHN. Furthermore, we perform some preliminary analyses of our distilled datasets to shed light on how they differ from naturally occurring data.
NeurIPS 2021. Code and datasets available at https://github.com/google-research/google-research/tree/master/kip
References in corpus (19)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Language Models are Few-Shot Learners
- Scaling Laws for Neural Language Models
- What makes ImageNet good for transfer learning?
- Explaining Neural Scaling Laws
- Coresets via Bilevel Optimization for Continual Learning and Streaming
- Enhanced Convolutional Neural Tangent Kernels
- The large learning rate phase of deep learning: the catapult mechanism
- Dataset Condensation with Differentiable Siamese Augmentation
- Neural Architecture Search on ImageNet in Four GPU Hours: A Theoretically Inspired Perspective
- What shapes feature representations? Exploring datasets, architectures, and training
- Neural Kernels Without Tangents
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of Generalization
- Infinite attention: NNGP and NTK for deep attention networks
- Flexible Dataset Distillation: Learn Labels Instead of Images
- Towards NNGP-guided Neural Architecture Search
- Asymptotics of Wide Convolutional Neural Networks
- Approximation and Learning with Deep Convolutional Models: a Kernel Perspective