10 papers
SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling
Haoran Xu, Hongyu Wang, Yifei Gao +3
On-policy distillation (OPD) trains a student on its own trajectories with dense per-token supervision from a stronger teacher, and often outperforms off-policy distillation and st…
Visual Para-Thinker++: A Single-Policy Multi-Agent Framework for Visual Reasoning
Haoran Xu, Hongyu Wang, Yifei Gao +4
Visual reasoning requires integrating evidence distributed across regions, attributes, and relations, making single-chain reasoning prone to early perceptual commitment and halluci…
Federated Balanced Learning
Jiaze Li, Haoran Xu, Wanyi Wu +9
Federated learning is a paradigm of joint learning in which clients collaborate by sharing model parameters instead of data. However, in the non-iid setting, the global model exper…
Federated Joint Learning for Domain and Class Generalization
Haoran Xu, Jiaze Li, Jianzhong Ju +1
Efficient fine-tuning of visual-language models like CLIP has become crucial due to their large-scale parameter size and extensive pretraining requirements. Existing methods typica…
Vision Also You Need: Navigating Out-of-Distribution Detection with Multimodal Large Language Model
Haoran Xu, Yanlin Liu, Zizhao Tong +8
Out-of-Distribution (OOD) detection is a critical task that has garnered significant attention. The emergence of CLIP has spurred extensive research into zero-shot OOD detection, o…
ImagebindDC: Compressing Multi-modal Data with Imagebind-based Condensation
Yue Min, Shaobo Wang, Jiaze Li +5
Data condensation techniques aim to synthesize a compact dataset from a larger one to enable efficient model training, yet while successful in unimodal settings, they often fail in…