3 papers
cs.CV2025
Detailed Object Description with Controllable Dimensions
Xinran Wang, Haiwen Zhang, Baoteng Li +5
Object description plays an important role for visually impaired individuals to understand and compare the differences between objects. Recent multimodal large language models(MLLM…
cs.CV2024
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
Xinran Wang, Muxi Diao, Baoteng Li +3
The Controllable Image Captioning Agent (CapAgent) is an innovative system designed to bridge the gap between user simplicity and professional-level outputs in image captioning tas…
cs.CV2024
Evaluating Attribute Comprehension in Large Vision-Language Models
Haiwen Zhang, Zixi Yang, Yuanzhi Liu +4
Currently, large vision-language models have gained promising progress on many downstream tasks. However, they still suffer many challenges in fine-grained visual understanding tas…