2 papers
cs.IR2026
SAGE: Semantic Attribute Graphs for Multi-Entity Visual Retrieval
Yongjoo Kim, Mincheol Kwon, Seonga Choi +5
Dense document images often contain many fine-grained visual and textual entities whose relevance depends on a user query. Standard vision-language retrievers encode cropped region…
cs.CV2026
Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding
Mincheol Kwon, Minseung Lee, Seonga Choi +7
Large Vision-Language Models (LVLMs) have shown strong performance across various multimodal tasks by leveraging the reasoning capabilities of Large Language Models (LLMs). However…