74 citations · 179 across the 43 of their papers we have counts for
7 papers · 2 filters
Injecting Semantic Concepts into End-to-End Image Captioning
Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu +5
Tremendous progress has been made in recent years in developing better image captioning models, yet most of them rely on a separate object detector to extract regional features. Re…
Semantically Distributed Robust Optimization for Vision-and-Language Inference
Tejas Gokhale, Abhishek Chaudhary, Pratyay Banerjee +2
Analysis of vision-and-language models has revealed their brittleness under linguistic phenomena such as paraphrasing, negation, textual entailment, and word substitutions with syn…
Weakly Supervised Relative Spatial Reasoning for Visual Question Answering
Pratyay Banerjee, Tejas Gokhale, Yezhou Yang +1
Vision-and-language (V\&L) reasoning necessitates perception of visual concepts such as objects and actions, understanding semantics and language grounding, and reasoning about the…
CLEVR_HYP: A Challenge Dataset and Baselines for Visual Question Answering with Hypothetical Actions over Images
Shailaja Keyur Sampat, Akshay Kumar, Yezhou Yang +1
Most existing research on visual question answering (VQA) is limited to information explicitly present in an image or a video. In this paper, we take visual understanding to a high…
Compressing Visual-linguistic Model via Knowledge Distillation
Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu +3
Despite exciting progress in pre-training for visual-linguistic (VL) representations, very few aspire to a small VL model. In this paper, we study knowledge distillation (KD) to ef…
Hierarchical and Partially Observable Goal-driven Policy Learning with Goals Relational Graph
Xin Ye, Yezhou Yang
We present a novel two-layer hierarchical reinforcement learning approach equipped with a Goals Relational Graph (GRG) for tackling the partially observable goal-driven task, such…