activity
20242026
collaborators

6 papers

cs.LG2026

Distill What the Student Can See: Fisher-Projected On-Policy Distillation for Vision-Language Models

Leyan Xue, Feng Xiong, Mingjun Ma +1

On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-token distributions at prefi…

cs.LG2026

MULTIBENCH++: A Unified and Comprehensive Multimodal Fusion Benchmarking Across Specialized Domains

Leyan Xue, Changqing Zhang, Kecheng Xue +3

Although multimodal fusion has made significant progress, its advancement is severely hindered by the lack of adequate evaluation benchmarks. Current fusion methods are typically e…

cs.LG2025

DOTA: Distributional Test-Time Adaptation of Vision-Language Models

Zongbo Han, Jialong Yang, Guangyu Wang +4

Vision-language foundation models (VLMs), such as CLIP, exhibit remarkable performance across a wide range of tasks. However, deploying these models can be unreliable when signific…

cs.CV2025

Retrieval-Augmented Prompt for OOD Detection

Ruisong Han, Zongbo Han, Jiahao Zhang +2

Out-of-Distribution (OOD) detection is crucial for the reliable deployment of machine learning models in-the-wild, enabling accurate identification of test samples that differ from…

cs.CV2025

Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach

Leyan Xue, Zongbo Han, Guangyu Wang +3

Vision-Language Models (VLMs) like CLIP achieve cross-modal semantic alignment through contrastive learning, exhibiting robust zero-shot generalization. Traditional prompt engineer…

cs.LG2024

Out-Of-Distribution Detection with Diversification (Provably)

Haiyun Yao, Zongbo Han, Huazhu Fu +3

Out-of-distribution (OOD) detection is crucial for ensuring reliable deployment of machine learning models. Recent advancements focus on utilizing easily accessible auxiliary outli…