8 citations · 10 across the 3 of their papers we have counts for
3 papers
cs.CV2026
Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models
Issa Sugiura, Keito Sasagawa, Keisuke Nakao +8
Developing vision-language models (VLMs) that generalize across diverse tasks requires large-scale training datasets with diverse content. In English, such datasets are typically c…
cs.CV2023★ 2 cited
SciCap+: A Knowledge Augmented Dataset to Study the Challenges of Scientific Figure Captioning
Zhishen Yang, Raj Dabre, Hideki Tanaka +1
In scholarly documents, figures provide a straightforward way of communicating scientific findings to readers. Automating figure caption generation helps move model understandings…
cs.CL2020★ 8 cited
Keyframe Segmentation and Positional Encoding for Video-guided Machine Translation Challenge 2020
Tosho Hirasawa, Zhishen Yang, Mamoru Komachi +1
Video-guided machine translation as one of multimodal neural machine translation tasks targeting on generating high-quality text translation by tangibly engaging both video and tex…