Towards Data-Free Domain Generalization
arXiv:2110.04545
Abstract
In this work, we investigate the unexplored intersection of domain generalization (DG) and data-free learning. In particular, we address the question: How can knowledge contained in models trained on different source domains be merged into a single model that generalizes well to unseen target domains, in the absence of source and target domain data? Machine learning models that can cope with domain shift are essential for real-world scenarios with often changing data distributions. Prior DG methods typically rely on using source domain data, making them unsuitable for private decentralized data. We define the novel problem of Data-Free Domain Generalization (DFDG), a practical setting where models trained on the source domains separately are available instead of the original datasets, and investigate how to effectively solve the domain generalization problem in that case. We propose DEKAN, an approach that extracts and fuses domain-specific knowledge from the available teacher models into a student model robust to domain shift. Our empirical evaluation demonstrates the effectiveness of our method which achieves first state-of-the-art results in DFDG by significantly outperforming data-free knowledge distillation and ensemble baselines.
Accepted at NeurIPS 2021 (DistShift Workshop) and ACML 2022
References in corpus (20)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Distilling the Knowledge in a Neural Network
- Improved Regularization of Convolutional Neural Networks with Cutout
- Deep Domain Confusion: Maximizing for Domain Invariance
- Domain Generalization via Invariant Feature Representation
- Model Adaptation: Unsupervised Domain Adaptation without Source Data
- Domain Generalization via Model-Agnostic Learning of Semantic Features
- Data-Free Knowledge Distillation for Deep Neural Networks
- Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy
- Improve Unsupervised Domain Adaptation with Mixup Training
- Heterogeneous Domain Generalization via Domain Mixup
- Zero-Shot Knowledge Distillation in Deep Networks
- In Search of Lost Domain Generalization
- Self-supervised Knowledge Distillation for Few-shot Learning
- Frustratingly Simple Domain Generalization via Image Stylization
- Domain Generalization with MixStyle
- Large-Scale Generative Data-Free Distillation
- SAND-mask: An Enhanced Gradient Masking Strategy for the Discovery of Invariances in Domain Generalization
- Gradient Matching for Domain Generalization
- Source-Free Adaptation to Measurement Shift via Bottom-Up Feature Restoration