activity
20182021
most citedReferring Segmentation in Images and Videos with Cross-Modal Self-Attention Network

88 citations · 193 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV202188 cited

Referring Segmentation in Images and Videos with Cross-Modal Self-Attention Network

Linwei Ye, Mrigank Rochan, Zhi Liu +2

We consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the…

cs.CV20204 cited

Adaptive Video Highlight Detection by Learning from User History

Mrigank Rochan, Mahesh Kumar Krishna Reddy, Linwei Ye +1

Recently, there is an increasing interest in highlight detection research where the goal is to create a short duration video from a longer video by extracting its interesting momen…

cs.CV20207 cited

Cross-Modal Weighting Network for RGB-D Salient Object Detection

Gongyang Li, Zhi Liu, Linwei Ye +2

Depth maps contain geometric clues for assisting Salient Object Detection (SOD). In this paper, we propose a novel Cross-Modal Weighting (CMW) strategy to encourage comprehensive i…

cs.CV202048 cited

Dual Convolutional LSTM Network for Referring Image Segmentation

Linwei Ye, Zhi Liu, Yang Wang

We consider referring image segmentation. It is a problem at the intersection of computer vision and natural language understanding. Given an input image and a referring expression…

cs.CV201946 cited

Cross-Modal Self-Attention Network for Referring Image Segmentation

Linwei Ye, Mrigank Rochan, Zhi Liu +1

We consider the problem of referring image segmentation. Given an input image and a natural language expression, the goal is to segment the object referred by the language expressi…

cs.CV2018

Video Summarization Using Fully Convolutional Sequence Networks

Mrigank Rochan, Linwei Ye, Yang Wang

This paper addresses the problem of video summarization. Given an input video, the goal is to select a subset of the frames to create a summary video that optimally captures the im…