activity
20162022
most citedAction Recognition with Joint Attention on Multi-Level Deep Features

13 citations · 22 across the 5 of their papers we have counts for

collaborators

10 papers

cs.CV20211 cited

Visual Question Answering based on Local-Scene-Aware Referring Expression Generation

Jung-Jun Kim, Dong-Gyu Lee, Jialin Wu +2

Visual question answering requires a deep understanding of both images and natural language. However, most methods mainly focus on visual concept; such as the relationships between…

cs.CV20208 cited

Improving VQA and its Explanations \\ by Comparing Competing Explanations

Jialin Wu, Liyan Chen, Raymond J. Mooney

Most recent state-of-the-art Visual Question Answering (VQA) systems are opaque black boxes that are only trained to fit the answer distribution given the question and visual conte…

cs.CV2019

Hidden State Guidance: Improving Image Captioning using An Image Conditioned Autoencoder

Jialin Wu, Raymond J. Mooney

Most RNN-based image captioning models receive supervision on the output words to mimic human captions. Therefore, the hidden states can only receive noisy gradient signals via lay…

cs.CV2019

Generating Question Relevant Captions to Aid Visual Question Answering

Jialin Wu, Zeyuan Hu, Raymond J. Mooney

Visual question answering (VQA) and image captioning require a shared body of general knowledge connecting language and vision. We present a novel approach to improve VQA performan…

cs.CV2019

Self-Critical Reasoning for Robust Visual Question Answering

Jialin Wu, Raymond J. Mooney

Visual Question Answering (VQA) deep-learning systems tend to capture superficial statistical correlations in the training data because of strong language priors and fail to genera…

cs.CV2018

Image Score: How to Select Useful Samples

Simiao Zuo, Jialin Wu

There has long been debates on how we could interpret neural networks and understand the decisions our models make. Specifically, why deep neural networks tend to be error-prone wh…