64 citations · 109 across the 8 of their papers we have counts for
10 papers
Contrastive Video-Language Segmentation
Chen Liang, Yawei Luo, Yu Wu +1
We focus on the problem of segmenting a certain object referred by a natural language sentence in video content, at the core of formulating a pinpoint vision-language relation. Whi…
Saying the Unseen: Video Descriptions via Dialog Agents
Ye Zhu, Yu Wu, Yi Yang +1
Current vision and language tasks usually take complete visual data (e.g., raw images or videos) as input, however, practical scenarios may often consist the situations where part…
Learning Audio-Visual Correlations from Variational Cross-Modal Generation
Ye Zhu, Yu Wu, Hugo Latapie +2
People can easily imagine the potential sound while seeing an event. This natural synchronization between audio and visual signals reveals their intrinsic correlations. To this end…
Learning to Anticipate Egocentric Actions by Imagination
Yu Wu, Linchao Zhu, Xiaohan Wang +2
Anticipating actions before they are executed is crucial for a wide range of practical applications, including autonomous driving and robotics. In this paper, we study the egocentr…
Describing Unseen Videos via Multi-Modal Cooperative Dialog Agents
Ye Zhu, Yu Wu, Yi Yang +1
With the arising concerns for the AI systems provided with direct access to abundant sensitive information, researchers seek to develop more reliable AI with implicit information s…
Unsupervised Person Re-identification via Softened Similarity Learning
Yutian Lin, Lingxi Xie, Yu Wu +2
Person re-identification (re-ID) is an important topic in computer vision. This paper studies the unsupervised setting of re-ID, which does not require any labeled information and…