2 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CL2024★ 1 cited
Empirical Analysis of Large Vision-Language Models against Goal Hijacking via Visual Prompt Injection
Subaru Kimura, Ryota Tanaka, Shumpei Miyawaki +2
We explore visual prompt injection (VPI) that maliciously exploits the ability of large vision-language models (LVLMs) to follow instructions drawn onto the input image. We propose…
cs.CL2023★ 2 cited
SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida +3
Visual question answering on document images that contain textual, visual, and layout information, called document VQA, has received much attention recently. Although many datasets…