20 citations · 43 across the 12 of their papers we have counts for
4 papers · 1 filter
Video-aided Unsupervised Grammar Induction
Songyang Zhang, Linfeng Song, Lifeng Jin +3
We investigate video-aided grammar induction, which learns a constituency parser from both unlabeled text and its corresponding video. Existing methods of multi-modal grammar induc…
Comprehensive Image Captioning via Scene Graph Decomposition
Yiwu Zhong, Liwei Wang, Jianshu Chen +2
We address the challenging problem of image captioning by revisiting the representation of image scene graph. At the core of our method lies the decomposition of a scene graph into…
Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
Liwei Wang, Jing Huang, Yin Li +3
Weakly supervised phrase grounding aims at learning region-phrase correspondences using only image-sentence pairs. A major challenge thus lies in the missing links between image re…
A Fast and Accurate One-Stage Approach to Visual Grounding
Zhengyuan Yang, Boqing Gong, Liwei Wang +3
We propose a simple, fast, and accurate one-stage approach to visual grounding, inspired by the following insight. The performances of existing propose-and-rank two-stage methods a…