1 paper
Hung-Jen Chen, Yu-Heng Ho, Ting-Yao Huang +4
Generative vision-language models (VLMs) such as Qwen-VL and LLaVA achieve strong zero-shot performance on tasks overlapping with their pretraining distribution, yet fail on specia…