2 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 1 cited
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
Ryota Tanaka, Taichi Iki, Kyosuke Nishida +2
We study the problem of completing various visual document understanding (VDU) tasks, e.g., question answering and information extraction, on real-world documents through human-wri…
cs.CL2023★ 2 cited
SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida +3
Visual question answering on document images that contain textual, visual, and layout information, called document VQA, has received much attention recently. Although many datasets…