2 citations · 6 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 2 cited
Large Language Models are Temporal and Causal Reasoners for Video Question Answering
Dohwan Ko, Ji Soo Lee, Wooyoung Kang +2
Large Language Models (LLMs) have shown remarkable performances on a wide range of natural language understanding and generation tasks. We observe that the LLMs provide effective p…
cs.CV2023★ 2 cited
Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models
Dohwan Ko, Ji Soo Lee, Miso Choi +3
Video Question Answering (VideoQA) is a challenging task that entails complex multi-modal reasoning. In contrast to multiple-choice VideoQA which aims to predict the answer given s…
cs.CV2023★ 2 cited
MELTR: Meta Loss Transformer for Learning to Fine-tune Video Foundation Models
Dohwan Ko, Joonmyung Choi, Hyeong Kyu Choi +3
Foundation models have shown outstanding performance and generalization capabilities across domains. Since most studies on foundation models mainly focus on the pretraining phase,…