1 paper · 1 filter
Kazuki Hayashi, Kazuma Onishi, Toma Suzuki +7
Large-scale Vision-Language Models (LVLMs) process both images and text, excelling in multimodal tasks such as image captioning and description generation. However, while these mod…