5 papers
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
Hao Dong, Hongzhao Li, Shupan Li +3
Despite the growing popularity of Multimodal Domain Generalization (MMDG) for enhancing model robustness, it remains unclear whether reported performance gains reflect genuine algo…
Balancing Multimodal Domain Generalization via Gradient Modulation and Projection
Hongzhao Li, Guohao Shen, Shupan Li +2
Multimodal Domain Generalization (MMDG) leverages the complementary strengths of multiple modalities to enhance model generalization on unseen domains. A central challenge in multi…
Towards Multimodal Domain Generalization with Few Labels
Hongzhao Li, Hao Dong, Hualei Wan +3
Multimodal models ideally should generalize to unseen domains while remaining data-efficient to reduce annotation costs. To this end, we introduce and study a new problem, Semi-Sup…
Compositional Zero-Shot Learning: A Survey
Ans Munir, Faisal Z. Qureshi, Mohsen Ali +1
Compositional Zero-Shot Learning (CZSL) is a critical task in computer vision that enables models to recognize unseen combinations of known attributes and objects during inference,…
TLAC: Two-stage LMM Augmented CLIP for Zero-Shot Classification
Ans Munir, Faisal Z. Qureshi, Muhammad Haris Khan +1
Contrastive Language-Image Pretraining (CLIP) has shown impressive zero-shot performance on image classification. However, state-of-the-art methods often rely on fine-tuning techni…