1 paper
Raehyuk Jung, Seungjun Yu, Hyunjung Shim
Vision-Language Models (VLMs) combine a vision encoder and a large language model (LLM) through alignment training, showing strong performance on multimodal tasks. A central compon…