2 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2025
Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
Ji Soo Lee, Byungoh Ko, Jaewon Cho +3
In text-video retrieval, auxiliary captions are often used to enhance video understanding, bridging the gap between the modalities. While recent advances in multi-modal large langu…
cs.CV2023★ 2 cited
Large Language Models are Temporal and Causal Reasoners for Video Question Answering
Dohwan Ko, Ji Soo Lee, Wooyoung Kang +2
Large Language Models (LLMs) have shown remarkable performances on a wide range of natural language understanding and generation tasks. We observe that the LLMs provide effective p…
cs.CV2023★ 2 cited
Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models
Dohwan Ko, Ji Soo Lee, Miso Choi +3
Video Question Answering (VideoQA) is a challenging task that entails complex multi-modal reasoning. In contrast to multiple-choice VideoQA which aims to predict the answer given s…