2 papers
cs.CV2025
Beyond Intermediate States: Explaining Visual Redundancy through Language
Dingchen Yang, Bowen Cao, Anran Zhang +3
Multi-modal Large Langue Models (MLLMs) often process thousands of visual tokens, which consume a significant portion of the context window and impose a substantial computational b…
cs.CV2024
Pensieve: Retrospect-then-Compare Mitigates Visual Hallucination
Dingchen Yang, Bowen Cao, Guang Chen +1
Multi-modal Large Language Models (MLLMs) demonstrate remarkable success across various vision-language tasks. However, they suffer from visual hallucination, where the generated r…