3 papers
cs.CL2026
Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering
Jongha Kim, Byungoh Ko, Jeehye Na +2
Despite the remarkable capabilities of Large Vision Language Models (LVLMs), they still lack detailed knowledge about specific entities. Retrieval-augmented Generation (RAG) is a w…
cs.CV2025
TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
Jongha Kim, Minseong Bae, Sanghyeok Lee +2
Table images present unique challenges for effective and efficient understanding due to the need for question-specific focus and the presence of redundant background regions. Exist…
cs.CV2025
VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning
Ji Soo Lee, Jongha Kim, Jeehye Na +2
Despite the advancements of Video Large Language Models (VideoLLMs) in various tasks, they struggle with fine-grained temporal understanding, such as Dense Video Captioning (DVC).…