3 papers
cs.CV2026
Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning
Mengzhao Wang, Yanli Ji, Wangmeng Zuo +2
Interleaved multimodal reasoning improves visual grounding by revisiting visual evidence during multi-step generation, yet existing methods typically rely on token replay, repeated…
cs.CV2024
Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding
Mengzhao Wang, Huafeng Li, Yafei Zhang +3
Video Paragraph Grounding (VPG) aims to precisely locate the most appropriate moments within a video that are relevant to a given textual paragraph query. However, existing methods…
cs.CV2024
Phrase Decoupling Cross-Modal Hierarchical Matching and Progressive Position Correction for Visual Grounding
Minghong Xie, Mengzhao Wang, Huafeng Li +3
Visual grounding has attracted wide attention thanks to its broad application in various visual language tasks. Although visual grounding has made significant research progress, ex…