Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
Meng Shen, Minghao Wu, Deepu Rajan
Object hallucination is a significant challenge that hinders the application of large vision-language models (LVLMs) in practice. We hypothesize that one possible origin of halluci…
cs.CV2024
GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation
Zhanyu Wang, Longyue Wang, Zhen Zhao +7
While the recent advances in Multimodal Large Language Models (MLLMs) constitute a significant leap forward in the field, these models are predominantly confined to the realm of in…