228 citations · 373 across the 14 of their papers we have counts for
5 papers · 2 filters
Simple is not Easy: A Simple Strong Baseline for TextVQA and TextCaps
Qi Zhu, Chenyu Gao, Peng Wang +1
Texts appearing in daily scenes that can be recognized by OCR (Optical Character Recognition) tools contain significant information, such as street name, product brand and prices.…
Hyperspectral Classification Based on Lightweight 3-D-CNN With Transfer Learning
Haokui Zhang, Ying Li, Yenan Jiang +3
Recently, hyperspectral image (HSI) classification approaches based on deep learning (DL) models have been proposed and shown promising performance. However, because of very limite…
Give Me Something to Eat: Referring Expression Comprehension with Commonsense Knowledge
Peng Wang, Dongyang Liu, Hui Li +1
Conventional referring expression comprehension (REF) assumes people to query something from an image by describing its visual appearance and spatial location, but in practice, we…
Structured Multimodal Attentions for TextVQA
Chenyu Gao, Qi Zhu, Peng Wang +4
In this paper, we propose an end-to-end structured multimodal attention (SMA) neural network to mainly solve the first two issues above. SMA first uses a structural graph represent…
Say As You Wish: Fine-grained Control of Image Caption Generation with Abstract Scene Graphs
Shizhe Chen, Qin Jin, Peng Wang +1
Humans are able to describe image contents with coarse to fine details as they wish. However, most image captioning models are intention-agnostic which can not generate diverse des…