3 papers
cs.CV2026
Synth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images
Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara
The ability of Large Vision Language Models (LVLMs) to read text within document images is crucial, as it enables various applications such as Document Visual Question Answering. T…
cs.CV2025
Evaluating Multimodal Large Language Models on Vertically Written Japanese Text
Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara
Multimodal Large Language Models (MLLMs) have seen rapid advances in recent years and are now being applied to visual document understanding tasks. They are expected to process a w…
cs.CL2024
Constructing Multimodal Datasets from Scratch for Rapid Development of a Japanese Visual Language Model
Keito Sasagawa, Koki Maeda, Issa Sugiura +3
To develop high-performing Visual Language Models (VLMs), it is essential to prepare multimodal resources, such as image-text pairs, interleaved data, and instruction data. While m…