3 papers
cs.AI2026
AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes
Yibo Wang, Rui Yang, Jisheng Dang +7
Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet affective reasoning remains vulnerable to shortcut behavior. Models may pred…
cs.CV2026
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
Xiang Fang, Wanlong Fang, Wei Ji +1
Video-language models are pivotal for tasks such as moment retrieval and highlight detection, yet they often struggle to capture the dynamic, non-linear interactions between tempor…
cs.CV2024
Effective and Efficient Adversarial Detection for Vision-Language Models via A Single Vector
Youcheng Huang, Fengbin Zhu, Jingkun Tang +4
Visual Language Models (VLMs) are vulnerable to adversarial attacks, especially those from adversarial images, which is however under-explored in literature. To facilitate research…