6 papers
Distill What the Student Can See: Fisher-Projected On-Policy Distillation for Vision-Language Models
Leyan Xue, Feng Xiong, Mingjun Ma +1
On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-token distributions at prefi…
MULTIBENCH++: A Unified and Comprehensive Multimodal Fusion Benchmarking Across Specialized Domains
Leyan Xue, Changqing Zhang, Kecheng Xue +3
Although multimodal fusion has made significant progress, its advancement is severely hindered by the lack of adequate evaluation benchmarks. Current fusion methods are typically e…
DOTA: Distributional Test-Time Adaptation of Vision-Language Models
Zongbo Han, Jialong Yang, Guangyu Wang +4
Vision-language foundation models (VLMs), such as CLIP, exhibit remarkable performance across a wide range of tasks. However, deploying these models can be unreliable when signific…
Retrieval-Augmented Prompt for OOD Detection
Ruisong Han, Zongbo Han, Jiahao Zhang +2
Out-of-Distribution (OOD) detection is crucial for the reliable deployment of machine learning models in-the-wild, enabling accurate identification of test samples that differ from…
Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach
Leyan Xue, Zongbo Han, Guangyu Wang +3
Vision-Language Models (VLMs) like CLIP achieve cross-modal semantic alignment through contrastive learning, exhibiting robust zero-shot generalization. Traditional prompt engineer…
Out-Of-Distribution Detection with Diversification (Provably)
Haiyun Yao, Zongbo Han, Huazhu Fu +3
Out-of-distribution (OOD) detection is crucial for ensuring reliable deployment of machine learning models. Recent advancements focus on utilizing easily accessible auxiliary outli…