2 papers
cs.CV2025
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
Minji Kim, Taekyung Kim, Bohyung Han
Video Large Language Models (VideoLLMs) extend the capabilities of vision-language models to spatiotemporal inputs, enabling tasks such as video question answering (VideoQA). Despi…
cs.CV2024
Leveraging Temporal Contextualization for Video Action Recognition
Minji Kim, Dongyoon Han, Taekyung Kim +1
We propose a novel framework for video understanding, called Temporally Contextualized CLIP (TC-CLIP), which leverages essential temporal information through global interactions in…