1 paper
Moran Yanuka, Assaf Ben Kish, Yonatan Bitton +2
Recent research increasingly focuses on training vision-language models (VLMs) with long, detailed image captions. However, small-scale VLMs often struggle to balance the richness…