2 papers
cs.CV2026
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
Xiang Fang, Wanlong Fang, Wei Ji +1
Video-language models are pivotal for tasks such as moment retrieval and highlight detection, yet they often struggle to capture the dynamic, non-linear interactions between tempor…
cs.CV2024
Effective and Efficient Adversarial Detection for Vision-Language Models via A Single Vector
Youcheng Huang, Fengbin Zhu, Jingkun Tang +4
Visual Language Models (VLMs) are vulnerable to adversarial attacks, especially those from adversarial images, which is however under-explored in literature. To facilitate research…