5 papers
SpatiO: Adaptive Test-Time Orchestration of Vision-Language Agents for Spatial Reasoning
Chan Yeong Hwang, Miso Choi, Sunghyun On +2
Understanding visual scenes requires not only recognizing objects but also reasoning about their spatial relationships. Unlike general vision-language tasks, spatial reasoning requ…
The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages
Miso Choi, Seonga Choi, Mincheol Kwon +3
Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs, forming distinct model lineages. It rema…
Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding
Mincheol Kwon, Minseung Lee, Seonga Choi +7
Large Vision-Language Models (LVLMs) have shown strong performance across various multimodal tasks by leveraging the reasoning capabilities of Large Language Models (LLMs). However…
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
Jihwan Park, Taehoon Song, Sanghyeok Lee +2
Vision-Language Models (VLMs) have been widely used in various visual recognition tasks due to their remarkable generalization capabilities. As these models grow in size and comple…
ReCo: Reminder Composition Mitigates Hallucinations in Vision-Language Models
Sotirios Panagiotis Chytas, Miso Choi, Hyunwoo J. Kim +1
Vision Language Models (VLMs) show impressive capabilities in integrating and reasoning with both visual and language data. But these models make mistakes. A common finding -- simi…