7 papers · 1 filter
Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization
Paul Hyunbin Cho, Jinhyuk Jang, SeokYoung Lee +7
Diffusion-based lip synchronization models achieve strong visual quality and audio-visual alignment, but full-sequence bidirectional attention and many denoising steps make them im…
WorldKV: Efficient World Memory with World Retrieval and Compression
Jung Yi, Minjae Kim, Paul Hyunbin Cho +3
Autoregressive video diffusion models have enabled real-time, action-conditioned world generation. However, sustaining a persistent world, where revisiting a previously seen viewpo…
Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression
Jung Yi, Wooseok Jang, Paul Hyunbin Cho +3
Recent advances in autoregressive video diffusion have enabled real-time frame streaming, yet existing solutions still suffer from temporal repetition, drift, and motion decelerati…
Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking
Soowon Son, Honggyu An, Jisu Nam +7
Despite achieving strong results on standard benchmarks, current point tracking methods rely on feature backbones that are rarely designed with the temporal coherence needed for ro…
MV-TAP: Tracking Any Point in Multi-View Videos
Jahyeok Koo, Inès Hyeonsu Kim, Mungyeom Kim +6
Multi-view camera systems enable rich observations of complex real-world scenes, and understanding dynamic objects in multi-view settings has become central to various applications…
Exploring Temporally-Aware Features for Point Tracking
Inès Hyeonsu Kim, Seokju Cho, Jiahui Huang +3
Point tracking in videos is a fundamental task with applications in robotics, video editing, and more. While many vision tasks benefit from pre-trained feature backbones to improve…