activity
20192021
most citedEnd-to-end Multi-modal Video Temporal Grounding

21 citations · 39 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV20212 cited

Video Salient Object Detection via Contrastive Features and Attention Modules

Yi-Wen Chen, Xiaojie Jin, Xiaohui Shen +1

Video salient object detection aims to find the most visually distinctive objects in a video. To explore the temporal dependencies, existing methods usually resort to recurrent neu…

cs.CV202121 cited

End-to-end Multi-modal Video Temporal Grounding

Yi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan Yang

We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from…

cs.CV20211 cited

Understanding Synonymous Referring Expressions via Contrastive Features

Yi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan Yang

Referring expression comprehension aims to localize objects identified by natural language descriptions. This is a challenging task as it requires understanding of both visual and…

cs.CV20206 cited

Regularizing Meta-Learning via Gradient Dropout

Hung-Yu Tseng, Yi-Wen Chen, Yi-Hsuan Tsai +3

With the growing attention on learning-to-learn new tasks using only a few examples, meta-learning has been widely used in numerous problems such as few-shot classification, reinfo…

cs.CV20198 cited

Referring Expression Object Segmentation with Caption-Aware Consistency

Yi-Wen Chen, Yi-Hsuan Tsai, Tiantian Wang +2

Referring expressions are natural language descriptions that identify a particular object within a scene and are widely used in our daily conversations. In this work, we focus on s…

cs.CV20191 cited

Unseen Object Segmentation in Videos via Transferable Representations

Yi-Wen Chen, Yi-Hsuan Tsai, Chu-Ya Yang +2

In order to learn object segmentation models in videos, conventional methods require a large amount of pixel-wise ground truth annotations. However, collecting such supervised data…