Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
CaptionQA: Is Your Caption as Useful as the Image Itself?
Shijia Yang, Yunong Liu, Bohan Zhai +5
Image captions serve as efficient surrogates for visual content in multimodal systems such as retrieval, recommendation, and multi-step agentic inference pipelines. Yet current eva…
cs.CV2024
HallE-Control: Controlling Object Hallucination in Large Multimodal Models
Bohan Zhai, Shijia Yang, Chenfeng Xu +4
Current Large Multimodal Models (LMMs) achieve remarkable progress, yet there remains significant uncertainty regarding their ability to accurately apprehend visual details, that i…