78 citations · 90 across the 9 of their papers we have counts for
17 papers
Reasoning with Multi-Structure Commonsense Knowledge in Visual Dialog
Shunyu Zhang, Xiaoze Jiang, Zequn Yang +2
Visual Dialog requires an agent to engage in a conversation with humans grounded in an image. Many studies on Visual Dialog focus on the understanding of the dialog history or the…
KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual Dialogue
Xiaoze Jiang, Siyi Du, Zengchang Qin +2
Visual dialogue is a challenging task that needs to extract implicit information from both visual (image) and textual (dialogue history) contexts. Classical approaches pay more att…
DAM: Deliberation, Abandon and Memory Networks for Generating Detailed and Non-repetitive Responses in Visual Dialogue
Xiaoze Jiang, Jing Yu, Yajing Sun +4
Visual Dialogue task requires an agent to be engaged in a conversation with human about an image. The ability of generating detailed and non-repetitive responses is crucial for the…
Multi-Level Network for High-Speed Multi-Person Pose Estimation
Ying Huang, Jiankai Zhuang, Zengchang Qin
In multi-person pose estimation, the left/right joint type discrimination is always a hard problem because of the similar appearance. Traditionally, we solve this problem by stacki…
FollowMeUp Sports: New Benchmark for 2D Human Keypoint Recognition
Ying Huang, Bin Sun, Haipeng Kan +2
Human pose estimation has made significant advancement in recent years. However, the existing datasets are limited in their coverage of pose variety. In this paper, we introduce a…
DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue
Xiaoze Jiang, Jing Yu, Zengchang Qin +4
Different from Visual Question Answering task that requires to answer only one question about an image, Visual Dialogue involves multiple questions which cover a broad range of vis…