2 citations · 9 across the 18 of their papers we have counts for
7 papers · 1 filter
Whether you can locate or not? Interactive Referring Expression Generation
Fulong Ye, Yuxing Long, Fangxiang Feng +1
Referring Expression Generation (REG) aims to generate unambiguous Referring Expressions (REs) for objects in a visual scene, with a dual task of Referring Expression Comprehension…
Towards Unifying Reference Expression Generation and Comprehension
Duo Zheng, Tao Kong, Ya Jing +2
Reference Expression Generation (REG) and Comprehension (REC) are two highly correlated tasks. Modeling REG and REC simultaneously for utilizing the relation between them is a prom…
Question-Driven Graph Fusion Network For Visual Question Answering
Yuxi Qian, Yuncong Hu, Ruonan Wang +2
Existing Visual Question Answering (VQA) models have explored various visual relationships between objects in the image to answer complex questions, which inevitably introduces irr…
Spot the Difference: A Cooperative Object-Referring Game in Non-Perfectly Co-Observable Scene
Duo Zheng, Fandong Meng, Qingyi Si +5
Visual dialog has witnessed great progress after introducing various vision-oriented goals into the conversation, especially such as GuessWhich and GuessWhat, where the only image…
Guessing State Tracking for Visual Dialogue
Wei Pang, Xiaojie Wang
The Guesser is a task of visual grounding in GuessWhat?! like visual dialogue. It locates the target object in an image supposed by an Oracle oneself over a question-answer based d…
Visual Dialogue State Tracking for Question Generation
Wei Pang, Xiaojie Wang
GuessWhat?! is a visual dialogue task between a guesser and an oracle. The guesser aims to locate an object supposed by the oracle oneself in an image by asking a sequence of Yes/N…