5 papers
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
Minjoon Jung, Junbin Xiao, Junghyun Kim +2
Do Video-LLMs have consistent temporal understanding when videos capture the same event from different viewpoints? To study this question, we introduce EgoExo-Con(sistency), a benc…
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
Minjoon Jung, Byoung-Tak Zhang, Lorenzo Torresani
Video temporal grounding (VTG) takes an untrimmed video and a natural-language query as input and localizes the temporal moment that best matches the query. Existing methods rely o…
Exploring Ordinal Bias in Action Recognition for Instructional Videos
Joochan Kim, Minjoon Jung, Byoung-Tak Zhang
Action recognition models have achieved promising results in understanding instructional videos. However, they often rely on dominant, dataset-specific action sequences rather than…
Confidence-guided Refinement Reasoning for Zero-shot Question Answering
Youwon Jang, Woo Suk Choi, Minjoon Jung +2
We propose Confidence-guided Refinement Reasoning (C2R), a novel training-free framework applicable to question-answering (QA) tasks across text, image, and video domains. C2R stra…
On the Consistency of Video Large Language Models in Temporal Comprehension
Minjoon Jung, Junbin Xiao, Byoung-Tak Zhang +1
Video large language models (Video-LLMs) can temporally ground language queries and retrieve video moments. Yet, such temporal comprehension capabilities are neither well-studied n…