From the 1 of 6 linked papers with an AI index.
6 papers
VGIF-Score: Interpretable and Diagnostic Evaluation of Spatio-Temporal Instruction Following in Video Generation
Songyu Xu, Xin Wang, Qiang Chen +6
The paper introduces VGIF-Score, an automated and interpretable framework that evaluates how well video generation models follow long, compositional instructions by parsing prompts…
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
Xinran Wang, Yuxuan Zhang, Xiao Zhang +7
Accurately detecting and localizing hallucinations is a critical task for ensuring high reliability of image captions. In the era of Multimodal Large Language Models (MLLMs), capti…
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
Xinran Wang, Muxi Diao, Yuanzhi Liu +4
Training text-to-image (T2I) models with detailed captions can significantly improve their generation quality. Existing methods often rely on simplistic metrics like caption length…
CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
Xinran Wang, Songyu Xu, Xiangxuan Shan +6
Cinematography is a cornerstone of film production and appreciation, shaping mood, emotion, and narrative through visual elements such as camera movement, shot composition, and lig…
Detailed Object Description with Controllable Dimensions
Xinran Wang, Haiwen Zhang, Baoteng Li +5
Object description plays an important role for visually impaired individuals to understand and compare the differences between objects. Recent multimodal large language models(MLLM…
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
Xinran Wang, Muxi Diao, Baoteng Li +3
The Controllable Image Captioning Agent (CapAgent) is an innovative system designed to bridge the gap between user simplicity and professional-level outputs in image captioning tas…