Rethinking Domain Generalization: Discriminability and Generalizability
arXiv:2309.16483 · doi:10.1109/TCSVT.2024.3422887
Abstract
Domain generalization(DG) endeavors to develop robust models that possess strong generalizability while preserving excellent discriminability. Nonetheless, pivotal DG techniques tend to improve the feature generalizability by learning domain-invariant representations, inadvertently overlooking the feature discriminability. On the one hand, the simultaneous attainment of generalizability and discriminability of features presents a complex challenge, often entailing inherent contradictions. This challenge becomes particularly pronounced when domain-invariant features manifest reduced discriminability owing to the inclusion of unstable factors, i.e., spurious correlations. On the other hand, prevailing domain-invariant methods can be categorized as category-level alignment, susceptible to discarding indispensable features possessing substantial generalizability and narrowing intra-class variations. To surmount these obstacles, we rethink DG from a new perspective that concurrently imbues features with formidable discriminability and robust generalizability, and present a novel framework, namely, Discriminative Microscopic Distribution Alignment~(DMDA). DMDA incorporates two core components: Selective Channel Pruning~(SCP) and Micro-level Distribution Alignment~(MDA). Concretely, SCP attempts to curtail redundancy within neural networks, prioritizing stable attributes conducive to accurate classification. This approach alleviates the adverse effect of spurious domain invariance and amplifies the feature discriminability. Besides, MDA accentuates micro-level alignment within each class, going beyond mere category-level alignment. Extensive experiments on four benchmark datasets corroborate that DMDA achieves comparable results to state-of-the-art methods in DG, underscoring the efficacy of our method.
Accepted to IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)
References in corpus (14)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation
- TransVOD: End-to-End Video Object Detection with Spatial-Temporal Transformers
- DMT: Dynamic Mutual Training for Semi-Supervised Learning
- Domain Generalization by Marginal Transfer Learning
- Context-Aware Mixup for Domain Adaptive Semantic Segmentation
- Improve Unsupervised Domain Adaptation with Mixup Training
- Instance-Aware Domain Generalization for Face Anti-Spoofing
- Uncertainty-Aware Consistency Regularization for Cross-Domain Semantic Segmentation
- Generative Domain Adaptation for Face Anti-Spoofing
- Simultaneous Semantic Alignment Network for Heterogeneous Domain Adaptation
- PIT: Position-Invariant Transform for Cross-FoV Domain Adaptation
- Domain Adaptive Semantic Segmentation via Regional Contrastive Consistency Regularization
- Rethinking Implicit Neural Representations for Vision Learners