2 papers
cs.CV2026
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
Meng Shen, Minghao Wu, Deepu Rajan
Object hallucination is a significant challenge that hinders the application of large vision-language models (LVLMs) in practice. We hypothesize that one possible origin of halluci…
cs.MM2024
Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning
Meng Shen, Yake Wei, Jianxiong Yin +3
Training multimodal models requires a large amount of labeled data. Active learning (AL) aim to reduce labeling costs. Most AL methods employ warm-start approaches, which rely on s…