2 papers
cs.CV2026
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
Xiang Fang, Wanlong Fang, Wei Ji +1
Video-language models are pivotal for tasks such as moment retrieval and highlight detection, yet they often struggle to capture the dynamic, non-linear interactions between tempor…
cs.CV2025
Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
Junyi Wang, Jinjiang Li, Guodong Fan +3
In the semantic segmentation of remote sensing images, acquiring complete ground objects is critical for achieving precise analysis. However, this task is severely hindered by two…