109 citations · 159 across the 5 of their papers we have counts for
7 papers
OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation
Jing Liu, Xinxin Zhu, Fei Liu +8
In this paper, we propose an Omni-perception Pre-Trainer (OPT) for cross-modal understanding and generation, by jointly modeling visual, text and audio resources. OPT is constructe…
CPTR: Full Transformer Network for Image Captioning
Wei Liu, Sihan Chen, Longteng Guo +2
In this paper, we consider the image captioning task from a new sequence-to-sequence prediction perspective and propose CaPtion TransformeR (CPTR) which takes the sequentialized ra…
Global-Local Propagation Network for RGB-D Semantic Segmentation
Sihan Chen, Xinxin Zhu, Wei Liu +2
Depth information matters in RGB-D semantic segmentation task for providing additional geometric information to color images. Most existing methods exploit a multi-stage fusion str…
Fast Sequence Generation with Multi-Agent Reinforcement Learning
Longteng Guo, Jing Liu, Xinxin Zhu +1
Autoregressive sequence Generation models have achieved state-of-the-art performance in areas like machine translation and image captioning. These models are autoregressive in that…
Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent Learning
Longteng Guo, Jing Liu, Xinxin Zhu +3
Most image captioning models are autoregressive, i.e. they generate each word by conditioning on previously generated words, which leads to heavy latency during inference. Recently…
Normalized and Geometry-Aware Self-Attention Network for Image Captioning
Longteng Guo, Jing Liu, Xinxin Zhu +3
Self-attention (SA) network has shown profound value in image captioning. In this paper, we improve SA from two aspects to promote the performance of image captioning. First, we pr…