2 citations · 4 across the 3 of their papers we have counts for
4 papers · 1 filter
Phrase Localization Without Paired Training Examples
Josiah Wang, Lucia Specia
Localizing phrases in images is an important part of image understanding and can be useful in many applications that require mappings between textual and visual information. Existi…
End-to-end Image Captioning Exploits Multimodal Distributional Similarity
Pranava Madhyastha, Josiah Wang, Lucia Specia
We hypothesize that end-to-end neural image captioning systems work seemingly well because they exploit and learn `distributional similarity' in a multimodal feature space by mappi…
Defoiling Foiled Image Captions
Pranava Madhyastha, Josiah Wang, Lucia Specia
We address the task of detecting foiled image captions, i.e. identifying whether a caption contains a word that has been deliberately replaced by a semantically similar word, thus…
Object Counts! Bringing Explicit Detections Back into Image Captioning
Josiah Wang, Pranava Madhyastha, Lucia Specia
The use of explicit object detectors as an intermediate step to image captioning - which used to constitute an essential stage in early work - is often bypassed in the currently do…