596 citations · 652 across the 18 of their papers we have counts for
9 papers · 1 filter
Scene Text Recognition with Semantics
Joshua Cesare Placidi, Yishu Miao, Zixu Wang +1
Scene Text Recognition (STR) models have achieved high performance in recent years on benchmark datasets where text images are presented with minimal noise. Traditional STR recogni…
Cross-Modal Generative Augmentation for Visual Question Answering
Zixu Wang, Yishu Miao, Lucia Specia
Data augmentation has been shown to effectively improve the performance of multimodal machine learning models. This paper introduces a generative model for data augmentation by lev…
Latent Variable Models for Visual Question Answering
Zixu Wang, Yishu Miao, Lucia Specia
Current work on Visual Question Answering (VQA) explore deterministic approaches conditioned on various types of image and question features. We posit that, in addition to image an…
MSVD-Turkish: A Comprehensive Multimodal Dataset for Integrated Vision and Language Research in Turkish
Begum Citamak, Ozan Caglayan, Menekse Kuyu +4
Automatic generation of video descriptions in natural language, also called video captioning, aims to understand the visual content of the video and produce a natural language sent…
Watch and Learn: Mapping Language and Noisy Real-world Videos with Self-supervision
Yujie Zhong, Linhai Xie, Sen Wang +2
In this paper, we teach machines to understand visuals and natural language by learning the mapping between sentences and noisy video snippets without explicit annotations. Firstly…
Phrase Localization Without Paired Training Examples
Josiah Wang, Lucia Specia
Localizing phrases in images is an important part of image understanding and can be useful in many applications that require mappings between textual and visual information. Existi…