1 paper
Amane Watahiki, Tomoki Doi, Taiga Shinozaki +4
One of the main objectives in developing large vision-language models (LVLMs) is to engineer systems that can assist humans with multimodal tasks, including interpreting descriptio…