6 papers
VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
Seohyun Lee, Seoung Choi, Dohwan Ko +2
As video corpora continue to expand in both scale and task complexity, there is increasing demand for approaches that retrieve relevant videos from large-scale corpora (inter-video…
DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning
Joonmyung Choi, Sanghyeok Lee, Jongha Kim +4
Recent advances in vision-language models have demonstrated remarkable performance across diverse multi-modal tasks, including document question answering that leverages structured…
MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models
Dohwan Ko, Jinyoung Park, Seoung Choi +3
Mixture-of-Experts (MoE) has emerged as an effective approach to reduce the computational overhead of Transformer architectures by sparsely activating a subset of parameters for ea…
Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval
Dohwan Ko, Ji Soo Lee, Minhyuk Choi +2
Text-Video Retrieval aims to find the most relevant text (or video) candidate given a video (or text) query from large-scale online databases. Recent work leverages multi-modal lar…
ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models
Dohwan Ko, Sihyeon Kim, Yumin Suh +4
Spatio-temporal reasoning is essential in understanding real-world environments in various fields, eg, autonomous driving and sports analytics. Recent advances have improved the sp…
LLaMo: Large Language Model-based Molecular Graph Assistant
Jinyoung Park, Minseong Bae, Dohwan Ko +1
Large Language Models (LLMs) have demonstrated remarkable generalization and instruction-following capabilities with instruction tuning. The advancements in LLMs and instruction tu…