most citedOPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation

21 citations · 23 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CV20211 cited

Exploiting Spatial-Temporal Semantic Consistency for Video Scene Parsing

Xingjian He, Weining Wang, Zhiyong Xu +3

Compared with image scene parsing, video scene parsing introduces temporal information, which can effectively improve the consistency and accuracy of prediction. In this paper, we…

cs.CV202121 cited

OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation

Jing Liu, Xinxin Zhu, Fei Liu +8

In this paper, we propose an Omni-perception Pre-Trainer (OPT) for cross-modal understanding and generation, by jointly modeling visual, text and audio resources. OPT is constructe…

cs.CV2021

Temporal Memory Attention for Video Semantic Segmentation

Hao Wang, Weining Wang, Jing Liu

Video semantic segmentation requires to utilize the complex temporal relations between frames of the video sequence. Previous works usually exploit accurate optical flow to leverag…

cs.CV20191 cited

Adaptive Context Network for Scene Parsing

Jun Fu, Jing Liu, Yuhang Wang +4

Recent works attempt to improve scene parsing performance by exploring different levels of contexts, and typically train a well-designed convolutional network to exploit useful con…

cs.CV2019

Vatex Video Captioning Challenge 2020: Multi-View Features and Hybrid Reward Strategies for Video Captioning

Xinxin Zhu, Longteng Guo, Peng Yao +3

This report describes our solution for the VATEX Captioning Challenge 2020, which requires generating descriptions for the videos in both English and Chinese languages. We identifi…

cs.CV2019

Aligning Linguistic Words and Visual Semantic Units for Image Captioning

Longteng Guo, Jing Liu, Jinhui Tang +3

Image captioning attempts to generate a sentence composed of several linguistic words, which are used to describe objects, attributes, and interactions in an image, denoted as visu…