6 papers · 1 filter
Be Consistent! Enhancing Robust Visual Reasoning in LVLMs with Consistency Constraints
Liqiang Jing, Xiong Zhou, Siddharth Varia +3
While Large Vision-Language Models (LVLMs) exhibit strong perceptual capabilities, they remain vulnerable in visual reasoning tasks. Existing benchmarks largely focus on symbolic m…
Multimodal Reference Visual Grounding
Yangxiao Lu, Ruosen Li, Liqiang Jing +5
Visual grounding focuses on detecting objects from images based on language expressions. Recent Large Vision-Language Models (LVLMs) have significantly advanced visual grounding pe…
FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
Liqiang Jing, Viet Lai, Seunghyun Yoon +2
Video Multimodal Large Language Models (VideoMLLMs) have achieved remarkable progress in both Video-to-Text and Text-to-Video tasks. However, they often suffer fro hallucinations,…
FIHA: Autonomous Hallucination Evaluation in Vision-Language Models with Davidson Scene Graphs
Bowen Yan, Zhengsong Zhang, Liqiang Jing +2
The rapid development of Large Vision-Language Models (LVLMs) often comes with widespread hallucination issues, making cost-effective and comprehensive assessments increasingly vit…
FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
Liqiang Jing, Xinya Du
Large Vision-Language Models (LVLMs) have demonstrated proficiency in tackling a variety of visual-language tasks. However, current LVLMs suffer from misalignment between text and…
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
Liqiang Jing, Guiming Hardy Chen, Ehsan Aghazadeh +2
Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in multimodal tasks, but visual object hallucination remains a persistent issue. It refers to scenarios whe…