12 papers
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
Hao Dong, Hongzhao Li, Shupan Li +3
Despite the growing popularity of Multimodal Domain Generalization (MMDG) for enhancing model robustness, it remains unclear whether reported performance gains reflect genuine algo…
Extremely Simple Multimodal Outlier Synthesis for Out-of-Distribution Detection and Segmentation
Moru Liu, Hao Dong, Jessica Kelly +2
Out-of-distribution (OOD) detection and segmentation are crucial for deploying machine learning models in safety-critical applications such as autonomous driving and robot-assisted…
Multimodal Learning for Arcing Detection in Pantograph-Catenary Systems
Hao Dong, Eleni Chatzi, Olga Fink
The pantograph-catenary interface is essential for ensuring uninterrupted and reliable power delivery in electrified rail systems. However, electrical arcing at this interface pose…
To Trust Or Not To Trust Your Vision-Language Model's Prediction
Hao Dong, Moru Liu, Jian Liang +2
Vision-Language Models (VLMs) have demonstrated strong capabilities in aligning visual and textual modalities, enabling a wide range of applications in multimodal understanding and…
Advances in Multimodal Adaptation and Generalization: From Traditional Approaches to Foundation Models
Hao Dong, Moru Liu, Kaiyang Zhou +4
In real-world scenarios, achieving domain adaptation and generalization poses significant challenges, as models must adapt to or generalize across unknown target distributions. Ext…
Adapting Vision-Language Models Without Labels: A Comprehensive Survey
Hao Dong, Lijun Sheng, Jian Liang +3
Vision-Language Models (VLMs) have demonstrated remarkable generalization capabilities across a wide range of tasks. However, their performance often remains suboptimal when direct…