7 papers
Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination
Zichuan Wang, Songlin Yang, Bo Peng +4
Large Vision-Language Models (LVLMs) often suffer from object hallucination, generating objects that are absent from the image. Prior work largely attributes this to insufficient v…
Endogenous Reprompting: Self-Evolving Cognitive Alignment for Unified Multimodal Models
Zhenchen Tang, Songlin Yang, Zichuan Wang +4
Unified Multimodal Models (UMMs) exhibit strong understanding, yet this capability often fails to effectively guide generation. We identify this as a Cognitive Gap: the model lacks…
Revisiting MLLM Based Image Quality Assessment: Errors and Remedy
Zhenchen Tang, Songlin Yang, Bo Peng +2
The rapid progress of multi-modal large language models (MLLMs) has boosted the task of image quality assessment (IQA). However, a key challenge arises from the inherent mismatch b…
HandEval: Taking the First Step Towards Hand Quality Evaluation in Generated Images
Zichuan Wang, Bo Peng, Songlin Yang +2
Although recent text-to-image (T2I) models have significantly improved the overall visual quality of generated images, they still struggle in the generation of accurate details in…
Towards Geometry-Grounded Dense Semantic Matching with VGGT Priors
Songlin Yang, Tianyi Wei, Yushi Lan +3
Semantic matching aims to establish pixel-level correspondences between instances of the same category and represents a fundamental task in computer vision. Existing approaches suf…
Instant Preference Alignment for Text-to-Image Diffusion Models
Yang Li, Songlin Yang, Xiaoxuan Han +4
Text-to-image (T2I) generation has greatly enhanced creative expression, yet achieving preference-aligned generation in a real-time and training-free manner remains challenging. Pr…