1 paper
Haoyi Zhou, Shuo Li, Tianyu Chen +3
While large vision-language models (VLMs) demonstrate strong long-context understanding, their prevalent small branches fail on linguistics-photography alignment for a limited wind…