activity
20192021
most citedCPTR: Full Transformer Network for Image Captioning

109 citations · 159 across the 5 of their papers we have counts for

collaborators

7 papers

cs.CV202121 cited

OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation

Jing Liu, Xinxin Zhu, Fei Liu +8

In this paper, we propose an Omni-perception Pre-Trainer (OPT) for cross-modal understanding and generation, by jointly modeling visual, text and audio resources. OPT is constructe…

cs.CV2021109 cited

CPTR: Full Transformer Network for Image Captioning

Wei Liu, Sihan Chen, Longteng Guo +2

In this paper, we consider the image captioning task from a new sequence-to-sequence prediction perspective and propose CaPtion TransformeR (CPTR) which takes the sequentialized ra…

cs.CV202111 cited

Global-Local Propagation Network for RGB-D Semantic Segmentation

Sihan Chen, Xinxin Zhu, Wei Liu +2

Depth information matters in RGB-D semantic segmentation task for providing additional geometric information to color images. Most existing methods exploit a multi-stage fusion str…

cs.CL20217 cited

Fast Sequence Generation with Multi-Agent Reinforcement Learning

Longteng Guo, Jing Liu, Xinxin Zhu +1

Autoregressive sequence Generation models have achieved state-of-the-art performance in areas like machine translation and image captioning. These models are autoregressive in that…

cs.CL202011 cited

Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent Learning

Longteng Guo, Jing Liu, Xinxin Zhu +3

Most image captioning models are autoregressive, i.e. they generate each word by conditioning on previously generated words, which leads to heavy latency during inference. Recently…

cs.CV2020

Normalized and Geometry-Aware Self-Attention Network for Image Captioning

Longteng Guo, Jing Liu, Xinxin Zhu +3

Self-attention (SA) network has shown profound value in image captioning. In this paper, we improve SA from two aspects to promote the performance of image captioning. First, we pr…