3 papers
cs.AI2026
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning
Zhiyuan Han, Beier Zhu, Wenwen Tong +6
We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning trajectories and (ii) exhibit u…
cs.LG2026
Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation
Sihan Wang, Xiyao Liu, Lianqing Liu +1
On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-level targets conditioned on a reference target. This works well…
cs.CV2025
Vision and Language Integration for Domain Generalization
Yanmei Wang, Xiyao Liu, Fupeng Chu +1
Domain generalization aims at training on source domains to uncover a domain-invariant feature space, allowing the model to perform robust generalization ability on unknown target…