6 papers
CS-CLIP: Compositional Scene Graph-guided CLIP for Robust Compositional Reasoning
SeongJun Jeong, Minjoon Jung, Woo Suk Choi +2
Vision-language models (VLMs) demonstrate strong performance across compositional reasoning benchmarks, which require reasoning over semantic perturbations of objects, attributes,…
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
Minjoon Jung, Byoung-Tak Zhang, Lorenzo Torresani
Video temporal grounding (VTG) takes an untrimmed video and a natural-language query as input and localizes the temporal moment that best matches the query. Existing methods rely o…
Kinematics-Driven Gaussian Shape Deformation for Blurry Monocular Dynamic Scenes
Yeon-Ji Song, Kiyoung Kwon, Junoh Lee +2
Reconstructing dynamic 3D scenes from blurry monocular videos is challenging as motion-induced blur entangles object motion and geometry, hindering geometric consistency. We presen…
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
Minjoon Jung, Junbin Xiao, Junghyun Kim +2
Do Video-LLMs have consistent temporal understanding when videos capture the same event from different viewpoints? To study this question, we introduce EgoExo-Con(sistency), a benc…
Exploring Ordinal Bias in Action Recognition for Instructional Videos
Joochan Kim, Minjoon Jung, Byoung-Tak Zhang
Action recognition models have achieved promising results in understanding instructional videos. However, they often rely on dominant, dataset-specific action sequences rather than…
On the Consistency of Video Large Language Models in Temporal Comprehension
Minjoon Jung, Junbin Xiao, Byoung-Tak Zhang +1
Video large language models (Video-LLMs) can temporally ground language queries and retrieve video moments. Yet, such temporal comprehension capabilities are neither well-studied n…