21 citations · 42 across the 9 of their papers we have counts for
3 papers · 2 filters
Video Salient Object Detection via Contrastive Features and Attention Modules
Yi-Wen Chen, Xiaojie Jin, Xiaohui Shen +1
Video salient object detection aims to find the most visually distinctive objects in a video. To explore the temporal dependencies, existing methods usually resort to recurrent neu…
End-to-end Multi-modal Video Temporal Grounding
Yi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan Yang
We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from…
Understanding Synonymous Referring Expressions via Contrastive Features
Yi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan Yang
Referring expression comprehension aims to localize objects identified by natural language descriptions. This is a challenging task as it requires understanding of both visual and…