3 papers
cs.IR2026
SAGE: Semantic Attribute Graphs for Multi-Entity Visual Retrieval
Yongjoo Kim, Mincheol Kwon, Seonga Choi +5
Dense document images often contain many fine-grained visual and textual entities whose relevance depends on a user query. Standard vision-language retrievers encode cropped region…
cs.CL2026
The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages
Miso Choi, Seonga Choi, Mincheol Kwon +3
Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs, forming distinct model lineages. It rema…
cs.CV2026
Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding
Mincheol Kwon, Minseung Lee, Seonga Choi +7
Large Vision-Language Models (LVLMs) have shown strong performance across various multimodal tasks by leveraging the reasoning capabilities of Large Language Models (LLMs). However…