1 paper
Young Kyun Jang, Junmo Kang, Yong Jae Lee +1
While advancements in Vision Language Models (VLMs) have significantly improved the alignment of visual and textual data, these models primarily focus on aligning images with short…