36 papers
Robustifying Vision-Language Models via Test-Time Prompt Adaptation
Xingyu Zhu, Huanshen Wu, Shuo Wang +4
Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing tes…
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
Sudong Wang, Weiquan Huang, Xiaomin Yu +9
The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followed by reinforcement learning with verifiab…
Class-frequency Guided Noise Schedule for Diffusion Models
Jiequan Cui, Beier Zhu, Qingshan Xu +3
In this paper, we are the first to examine the correlations between class frequency and the multi-scale noise schedule within diffusion models. For score-based generative models, l…
MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy
Zhiyuan Han, Beier Zhu, Wenwen Tong +8
We find that explicit reasoning does not necessarily translate into better multimodal emotion recognition (MER) accuracy, even though it makes predictions more interpretable. Speci…
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning
Zhiyuan Han, Beier Zhu, Wenwen Tong +6
We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning trajectories and (ii) exhibit u…
Detail++: Training-Free Detail Enhancer for T2I Diffusion Models
Lifeng Chen, Jiner Wang, Zihao Pan +3
Recent advances in text-to-image (T2I) generation have led to impressive visual results. However, these models still face significant challenges when handling complex prompt, parti…