2 papers
cs.CV2026
Fine-Grained Multi Image Object Hallucination Benchmark
Joonki Min, Chaeyun Kim, Hyungwook Choi +4
Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundam…
cs.CV2026
A More Word-like Image Tokenization for MLLMs
Hyun Lee, Hyemin Jeong, Yejin Kim +4
Modern multimodal large language models (MLLMs) typically keep the language model fixed and train a visual projector that maps the pixels into a sequence of tokens in its embedding…