activity
20182021
most citedAuto-Encoding Scene Graphs for Image Captioning

26 citations · 55 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CV20215 cited

Causal Attention for Vision-Language Tasks

Xu Yang, Hanwang Zhang, Guojun Qi +1

We present a novel attention mechanism: Causal Attention (CATT), to remove the ever-elusive confounding effect in existing attention-based vision-language models. This effect cause…

cs.CV20206 cited

Finding It at Another Side: A Viewpoint-Adapted Matching Encoder for Change Captioning

Xiangxi Shi, Xu Yang, Jiuxiang Gu +2

Change Captioning is a task that aims to describe the difference between images with natural language. Most existing methods treat this problem as a difference judgment without the…

cs.CV201918 cited

Learning to Collocate Neural Modules for Image Captioning

Xu Yang, Hanwang Zhang, Jianfei Cai

We do not speak word by word from scratch; our brain quickly structures a pattern like \textsc{sth do sth at someplace} and then fill in the detailed descriptions. To render existi…

cs.CV2019

Unpaired Image Captioning via Scene Graph Alignments

Jiuxiang Gu, Shafiq Joty, Jianfei Cai +3

Most of current image captioning models heavily rely on paired image-caption datasets. However, getting large scale image-caption paired data is labor-intensive and time-consuming.…

cs.CV201826 cited

Auto-Encoding Scene Graphs for Image Captioning

Xu Yang, Kaihua Tang, Hanwang Zhang +1

We propose Scene Graph Auto-Encoder (SGAE) that incorporates the language inductive bias into the encoder-decoder image captioning framework for more human-like captions. Intuitive…

cs.CV2018

Shuffle-Then-Assemble: Learning Object-Agnostic Visual Relationship Features

Xu Yang, Hanwang Zhang, Jianfei Cai

Due to the fact that it is prohibitively expensive to completely annotate visual relationships, i.e., the (obj1, rel, obj2) triplets, relationship models are inevitably biased to o…