11 citations · 14 across the 3 of their papers we have counts for
5 papers · 1 filter
Visual Reasoning: from State to Transformation
Xin Hong, Yanyan Lan, Liang Pang +2
Most existing visual reasoning tasks, such as CLEVR in VQA, ignore an important factor, i.e.~transformation. They are solely defined to test how well machines understand concepts a…
Visual Transformation Telling
Wanqing Cui, Xin Hong, Yanyan Lan +3
Humans can naturally reason from superficial state differences (e.g. ground wetness) to transformations descriptions (e.g. raining) according to their life experience. In this pape…
WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
Yuqi Huo, Manli Zhang, Guangzhen Liu +32
Multi-modal pre-training models have been intensively explored to bridge vision and language in recent years. However, most of them explicitly model the cross-modal interaction bet…
Transformation Driven Visual Reasoning
Xin Hong, Yanyan Lan, Liang Pang +2
This paper defines a new visual reasoning paradigm by introducing an important factor, i.e.~transformation. The motivation comes from the fact that most existing visual reasoning t…
Deep Fusion Network for Image Completion
Xin Hong, Pengfei Xiong, Renhe Ji +1
Deep image completion usually fails to harmonically blend the restored image into existing content, especially in the boundary area. This paper handles with this problem from a new…