10 papers
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
Minseok Kang, Minhyeok Lee, Jungho Lee +6
As Video Large Language Models (Video-LLMs) scale to longer and more complex videos, their inference cost grows rapidly due to the large volume of visual tokens accumulated across…
CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation
Inseok Jeon, Suhwan Cho, Minhyeok Lee +6
Recent advances in unsupervised video object segmentation have highlighted the potential of two-stream architectures that integrate appearance and motion cues. However, fully lever…
MoRGS: Efficient Per-Gaussian Motion Reasoning for Streamable Dynamic 3D Scenes
Wonjoon Lee, Sungmin Woo, Donghyeong Kim +3
Online reconstruction of dynamic scenes aims to learn from streaming multi-view inputs under low-latency constraints. The fast training and real-time rendering capabilities of 3D G…
Revisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning
Minseok Kang, Minhyeok Lee, Minjung Kim +5
Weakly-supervised video scene graph generation (WS-VSGG) aims to parse video content into structured relational triplets without bounding box annotations and with only sparse tempo…
Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal Grounding
Minseok Kang, Minhyeok Lee, Minjung Kim +2
Video Temporal Grounding (VTG) aims to localize temporal segments in long, untrimmed videos that align with a given natural language query. This task typically comprises two subtas…
DepthFlow: Exploiting Depth-Flow Structural Correlations for Unsupervised Video Object Segmentation
Suhwan Cho, Minhyeok Lee, Jungho Lee +2
Unsupervised video object segmentation (VOS) aims to detect the most prominent object in a video. Recently, two-stream approaches that leverage both RGB images and optical flow hav…