19 papers
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
Minseok Kang, Minhyeok Lee, Jungho Lee +6
As Video Large Language Models (Video-LLMs) scale to longer and more complex videos, their inference cost grows rapidly due to the large volume of visual tokens accumulated across…
Seen-to-Scene: Keep the Seen, Generate the Unseen for Video Outpainting
Inseok Jeon, Minhyeok Lee, Seunghoon Lee +3
Video outpainting aims to expand the visible content of a video beyond the original frame boundaries while preserving spatial fidelity and temporal coherence across frames. Existin…
CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation
Inseok Jeon, Suhwan Cho, Minhyeok Lee +6
Recent advances in unsupervised video object segmentation have highlighted the potential of two-stream architectures that integrate appearance and motion cues. However, fully lever…
Revisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning
Minseok Kang, Minhyeok Lee, Minjung Kim +5
Weakly-supervised video scene graph generation (WS-VSGG) aims to parse video content into structured relational triplets without bounding box annotations and with only sparse tempo…
SwiftVGGT: A Scalable Visual Geometry Grounded Transformer for Large-Scale Scenes
Jungho Lee, Minhyeok Lee, Sunghun Yang +2
3D reconstruction in large-scale scenes is a fundamental task in 3D perception, but the inherent trade-off between accuracy and computational efficiency remains a significant chall…
STATIC : Surface Temporal Affine for TIme Consistency in Video Monocular Depth Estimation
Sunghun Yang, Minhyeok Lee, Suhwan Cho +2
Video monocular depth estimation is essential for applications such as autonomous driving, AR/VR, and robotics. Recent transformer-based single-image monocular depth estimation mod…