7 papers
Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination
Zichuan Wang, Songlin Yang, Bo Peng +4
Large Vision-Language Models (LVLMs) often suffer from object hallucination, generating objects that are absent from the image. Prior work largely attributes this to insufficient v…
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
Songlin Yang, Haobin Zhong, Ruilin Zhang +23
The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such demanding quality, the community tr…
Endogenous Reprompting: Self-Evolving Cognitive Alignment for Unified Multimodal Models
Zhenchen Tang, Songlin Yang, Zichuan Wang +4
Unified Multimodal Models (UMMs) exhibit strong understanding, yet this capability often fails to effectively guide generation. We identify this as a Cognitive Gap: the model lacks…
Revisiting MLLM Based Image Quality Assessment: Errors and Remedy
Zhenchen Tang, Songlin Yang, Bo Peng +2
The rapid progress of multi-modal large language models (MLLMs) has boosted the task of image quality assessment (IQA). However, a key challenge arises from the inherent mismatch b…
HandEval: Taking the First Step Towards Hand Quality Evaluation in Generated Images
Zichuan Wang, Bo Peng, Songlin Yang +2
Although recent text-to-image (T2I) models have significantly improved the overall visual quality of generated images, they still struggle in the generation of accurate details in…
CLIP-AGIQA: Boosting the Performance of AI-Generated Image Quality Assessment with CLIP
Zhenchen Tang, Zichuan Wang, Bo Peng +1
With the rapid development of generative technologies, AI-Generated Images (AIGIs) have been widely applied in various aspects of daily life. However, due to the immaturity of the…