1 paper
Yuanzhi Liu, Shousheng Zhao, Bo Zhou +2
Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to temporal staleness, data conta…