collaborators

36 papers

cs.CV2026

Robustifying Vision-Language Models via Test-Time Prompt Adaptation

Xingyu Zhu, Huanshen Wu, Shuo Wang +4

Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing tes…

cs.CV2026

Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL

Sudong Wang, Weiquan Huang, Xiaomin Yu +9

The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followed by reinforcement learning with verifiab…

cs.LG2026

Class-frequency Guided Noise Schedule for Diffusion Models

Jiequan Cui, Beier Zhu, Qingshan Xu +3

In this paper, we are the first to examine the correlations between class frequency and the multi-scale noise schedule within diffusion models. For score-based generative models, l…

cs.AI2026

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy

Zhiyuan Han, Beier Zhu, Wenwen Tong +8

We find that explicit reasoning does not necessarily translate into better multimodal emotion recognition (MER) accuracy, even though it makes predictions more interpretable. Speci…

cs.AI2026

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning

Zhiyuan Han, Beier Zhu, Wenwen Tong +6

We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning trajectories and (ii) exhibit u…

cs.CV2026

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models

Lifeng Chen, Jiner Wang, Zihao Pan +3

Recent advances in text-to-image (T2I) generation have led to impressive visual results. However, these models still face significant challenges when handling complex prompt, parti…