15 citations · 43 across the 8 of their papers we have counts for
3 papers · 1 filter
Visual Recipe Flow: A Dataset for Learning Visual State Changes of Objects with Recipe Flows
Keisuke Shirai, Atsushi Hashimoto, Taichi Nishimura +4
We present a new multimodal dataset called Visual Recipe Flow, which enables us to learn each cooking action result in a recipe text. The dataset consists of object state changes a…
Removing Word-Level Spurious Alignment between Images and Pseudo-Captions in Unsupervised Image Captioning
Ukyo Honda, Yoshitaka Ushiku, Atsushi Hashimoto +2
Unsupervised image captioning is a challenging task that aims at generating captions without the supervision of image-sentence pairs, but only with images and sentences drawn from…
Customized Image Narrative Generation via Interactive Visual Question Generation and Answering
Andrew Shin, Yoshitaka Ushiku, Tatsuya Harada
Image description task has been invariably examined in a static manner with qualitative presumptions held to be universally applicable, regardless of the scope or target of the des…