21 citations · 36 across the 5 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
CaptionQA: Is Your Caption as Useful as the Image Itself?
Shijia Yang, Yunong Liu, Bohan Zhai +5
Image captions serve as efficient surrogates for visual content in multimodal systems such as retrieval, recommendation, and multi-step agentic inference pipelines. Yet current eva…
cs.CV2023
HallE-Control: Controlling Object Hallucination in Large Multimodal Models
Bohan Zhai, Shijia Yang, Chenfeng Xu +4
Current Large Multimodal Models (LMMs) achieve remarkable progress, yet there remains significant uncertainty regarding their ability to accurately apprehend visual details, that i…
cs.CV2022★ 13 cited
Multitask Vision-Language Prompt Tuning
Sheng Shen, Shijia Yang, Tianjun Zhang +4
Prompt Tuning, conditioning on task-specific learned prompt vectors, has emerged as a data-efficient and parameter-efficient method for adapting large pretrained vision-language mo…