15 papers
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
Xiang Fang, Wanlong Fang, Wei Ji +1
Video-language models are pivotal for tasks such as moment retrieval and highlight detection, yet they often struggle to capture the dynamic, non-linear interactions between tempor…
Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation
Xiang Fang, Wanlong Fang, Changshuo Wang
Vision-Language Navigation in Continuous Environments (VLN-CE) poses a formidable challenge for autonomous agents, requiring seamless integration of natural language instructions a…
Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition
Wanlong Fang, Tianle Zhang, Wen Tao +1
Understanding how multimodal large language models use different modalities is important for reliable reasoning. We employ Partial Information Decomposition (PID) as a decision-lev…
SLAP: The Semantic Least Action Principle for Variational Video-Language Modeling
Xiang Fang, Wanlong Fang
In the era of Large Video-Language Models (LVLMs), the computational necessity of sparse frame sampling creates a fundamental ``temporal gap'', rendering models blind to critical c…
Immuno-VLM: Immunizing Large Vision-Language Models via Generative Semantic Antibodies for Open-World Trustworthiness
Xiang Fang, Wanlong Fang, Wei Ji
Large Vision-Language Models have achieved unprecedented success in zero-shot recognition by aligning visual features with broad semantic concepts. However, this semantic abstracti…
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding
Xiang Fang, Daizong Liu, Wanlong Fang +4
This paper addresses the task of temporal sentence grounding (TSG). Although many respectable works have made decent achievements in this important topic, they severely rely on mas…