12 papers
Context Blindness in DPO: Mitigating Object Hallucination in MLLMs via Context-Calibrated Preference Optimization
Byungoh Ko, Jinyoung Park, Jongha Kim +3
Multimodal large language models (MLLMs) have made rapid progress, yet they still exhibit object hallucination, generating plausible but incorrect descriptions that are inconsisten…
VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
Seohyun Lee, Seoung Choi, Dohwan Ko +2
As video corpora continue to expand in both scale and task complexity, there is increasing demand for approaches that retrieve relevant videos from large-scale corpora (inter-video…
GLINT: Sparsely Gated Vision-Language Alignment for Fine-Grained Radiology Representations
Jonggwon Park, Seongeun Lee, Junhyun Park +6
Vision-language models (VLMs) for radiology have emerged as a scalable paradigm by leveraging image-report pairs naturally produced in clinical workflows. However, this pairing rev…
Retrieve What's Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation
Minseok Joo, Dogyun Park, Taehoon Lee +2
Maintaining long-term geometric consistency remains challenging for long-horizon autoregressive video generation. Memory-augmented generative models address this by retrieving hist…
DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning
Joonmyung Choi, Sanghyeok Lee, Jongha Kim +4
Recent advances in vision-language models have demonstrated remarkable performance across diverse multi-modal tasks, including document question answering that leverages structured…
RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection
Jihwan Park, Chanhyeong Yang, Jinyoung Park +2
Weakly-supervised Human-Object Interaction (HOI) detection is essential for scalable scene understanding, as it learns interactions from only image-level annotations. Due to the la…