9 papers
Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking
Soowon Son, Honggyu An, Jisu Nam +7
Despite achieving strong results on standard benchmarks, current point tracking methods rely on feature backbones that are rarely designed with the temporal coherence needed for ro…
Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization
Paul Hyunbin Cho, Jinhyuk Jang, SeokYoung Lee +7
Diffusion-based lip synchronization models achieve strong visual quality and audio-visual alignment, but full-sequence bidirectional attention and many denoising steps make them im…
MATRIX: Mask Track Alignment for Interaction-aware Video Generation
Siyoon Jin, Seongchan Kim, Dahyun Chung +5
Video DiTs have advanced video generation, yet they still struggle to model multi-instance or subject-object interactions. This raises a key question: How do these models internall…
WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation
Jisu Nam, Yicong Hong, Chun-Hao Paul Huang +9
Recent advances in video diffusion transformers have enabled interactive gaming world models that allow users to explore generated environments over extended horizons. However, exi…
Grounding World Simulation Models in a Real-World Metropolis
Junyoung Seo, Hyunwook Choi, Minkyung Kwon +10
What if a world simulation model could render not an imagined environment but a city that actually exists? Prior generative world models synthesize visually plausible yet artificia…
CORAL: Correspondence Alignment for Improved Virtual Try-On
Jiyoung Kim, Youngjin Shin, Siyoon Jin +6
Existing methods for Virtual Try-On (VTON) often struggle to preserve fine garment details, especially in unpaired settings where accurate person-garment correspondence is required…