activity
20192021
most citedCPTR: Full Transformer Network for Image Captioning

109 citations · 148 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CV202121 cited

OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation

Jing Liu, Xinxin Zhu, Fei Liu +8

In this paper, we propose an Omni-perception Pre-Trainer (OPT) for cross-modal understanding and generation, by jointly modeling visual, text and audio resources. OPT is constructe…

cs.CV2021109 cited

CPTR: Full Transformer Network for Image Captioning

Wei Liu, Sihan Chen, Longteng Guo +2

In this paper, we consider the image captioning task from a new sequence-to-sequence prediction perspective and propose CaPtion TransformeR (CPTR) which takes the sequentialized ra…

cs.CL20217 cited

Fast Sequence Generation with Multi-Agent Reinforcement Learning

Longteng Guo, Jing Liu, Xinxin Zhu +1

Autoregressive sequence Generation models have achieved state-of-the-art performance in areas like machine translation and image captioning. These models are autoregressive in that…

cs.CL202011 cited

Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent Learning

Longteng Guo, Jing Liu, Xinxin Zhu +3

Most image captioning models are autoregressive, i.e. they generate each word by conditioning on previously generated words, which leads to heavy latency during inference. Recently…

cs.CV2020

Normalized and Geometry-Aware Self-Attention Network for Image Captioning

Longteng Guo, Jing Liu, Xinxin Zhu +3

Self-attention (SA) network has shown profound value in image captioning. In this paper, we improve SA from two aspects to promote the performance of image captioning. First, we pr…

cs.CV2019

Vatex Video Captioning Challenge 2020: Multi-View Features and Hybrid Reward Strategies for Video Captioning

Xinxin Zhu, Longteng Guo, Peng Yao +3

This report describes our solution for the VATEX Captioning Challenge 2020, which requires generating descriptions for the videos in both English and Chinese languages. We identifi…