38 citations · 193 across the 40 of their papers we have counts for
5 papers · 1 filter
Visual Reasoning: from State to Transformation
Xin Hong, Yanyan Lan, Liang Pang +2
Most existing visual reasoning tasks, such as CLEVR in VQA, ignore an important factor, i.e.~transformation. They are solely defined to test how well machines understand concepts a…
Visual Transformation Telling
Wanqing Cui, Xin Hong, Yanyan Lan +3
Humans can naturally reason from superficial state differences (e.g. ground wetness) to transformations descriptions (e.g. raining) according to their life experience. In this pape…
Multi-video Moment Ranking with Multimodal Clue
Danyang Hou, Liang Pang, Yanyan Lan +2
Video corpus moment retrieval~(VCMR) is the task of retrieving a relevant video moment from a large corpus of untrimmed videos via a natural language query. State-of-the-art work f…
WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
Yuqi Huo, Manli Zhang, Guangzhen Liu +32
Multi-modal pre-training models have been intensively explored to bridge vision and language in recent years. However, most of them explicitly model the cross-modal interaction bet…
Transformation Driven Visual Reasoning
Xin Hong, Yanyan Lan, Liang Pang +2
This paper defines a new visual reasoning paradigm by introducing an important factor, i.e.~transformation. The motivation comes from the fact that most existing visual reasoning t…