2 papers
cs.CV2026
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
Yuanhong Zhang, Zhaoyang Wang, Xin Zhang +2
Large Vision-Language Models (LVLMs) have achieved remarkable success across cross-modal tasks but remain hindered by hallucinations, producing textual outputs inconsistent with vi…
cs.CV2024
Efficient Large Multi-modal Models via Visual Context Compression
Jieneng Chen, Luoxin Ye, Ju He +3
While significant advancements have been made in compressed representations for text embeddings in large language models (LLMs), the compression of visual tokens in multi-modal LLM…