13 papers
OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation
Shuimu Chen, Jing Jin, Nan Su +5
Large vision-language models (LVLMs) have recently shown strong potential for industrial anomaly detection (IAD) by providing image-level anomaly judgments and interpretable defect…
OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing
Zebang Cheng, Shuimu Chen, Boxue Yang +7
Reinforcement learning for multimodal large language models (MLLMs) is often hindered by severe reward sparsity in complex reasoning tasks. This challenge is particularly pronounce…
Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
Shuimu Chen, Yuteng Chen, Yuanshen Guan +7
Current multimodal reflection mechanisms for long video understanding predominantly rely on closed-loop self-reflection within internal parameters. Lacking objective external evide…
HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions
Yifei Dong, Fengyi Wu, Qi He +9
Vision-and-Language Navigation (VLN) has been studied mainly in either discrete or continuous spaces, with little attention to dynamic, crowded environments. We present HA-VLN 2.0,…
MER 2026: From Discriminative Emotion Recognition to Generative Emotion Understanding
Zheng Lian, Xiaojiang Peng, Kele Xu +15
MER2026 marks the fourth edition of the MER series of challenges. The MER series provides valuable data resources to the research community and offers tasks centered on recent rese…
EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs
He Hu, Tengjin Weng, Zebang Cheng +5
Recent multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and generation, and are increasingly used in applications such as social ro…