1 paper
Kosuke Sakurai, Tatsuya Ishii, Ryotaro Shimizu +2
In recent years, considerable research has been conducted on vision-language models that handle both image and text data; these models are being applied to diverse downstream tasks…