3 papers
cs.CV2026
MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models
Yuanzhi Liu, Shousheng Zhao, Bo Zhou +2
Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to temporal staleness, data conta…
cs.CV2025
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
Xinran Wang, Muxi Diao, Yuanzhi Liu +4
Training text-to-image (T2I) models with detailed captions can significantly improve their generation quality. Existing methods often rely on simplistic metrics like caption length…
cs.CV2024
Evaluating Attribute Comprehension in Large Vision-Language Models
Haiwen Zhang, Zixi Yang, Yuanzhi Liu +4
Currently, large vision-language models have gained promising progress on many downstream tasks. However, they still suffer many challenges in fine-grained visual understanding tas…