6 papers
GimbalDiffusion: Gravity-Aware Camera Control for Video Generation
Frédéric Fortier-Chouinard, Yannick Hold-Geoffroy, Valentin Deschaintre +2
Recent progress in text-to-video generation has achieved remarkable realism, yet fine-grained control over camera motion and orientation remains elusive, especially with extreme tr…
EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation
Jente Vandersanden, Matheus Gadelha, Chun-Hao P. Huang +2
Multi-shot video generation requires maintaining a consistent appearance of recurring entities across shots while remaining faithful to shot-specific text prompts. Recent autoregre…
Frame In-N-Out: Unbounded Controllable Image-to-Video Generation
Boyang Wang, Xuweiyi Chen, Matheus Gadelha +1
Controllability, temporal coherence, and detail synthesis remain the most critical challenges in video generation. In this paper, we focus on a commonly used yet underexplored cine…
3D-Fixup: Advancing Photo Editing with 3D Priors
Yen-Chi Cheng, Krishna Kumar Singh, Jae Shin Yoon +5
Despite significant advances in modeling image priors via diffusion models, 3D-aware image editing remains challenging, in part because the object is only specified via a single im…
PreciseCam: Precise Camera Control for Text-to-Image Generation
Edurne Bernal-Berdun, Ana Serrano, Belen Masia +4
Images as an artistic medium often rely on specific camera angles and lens distortions to convey ideas or emotions; however, such precise control is missing in current text-to-imag…
Motion Modes: What Could Happen Next?
Karran Pandey, Matheus Gadelha, Yannick Hold-Geoffroy +3
Predicting diverse object motions from a single static image remains challenging, as current video generation models often entangle object movement with camera motion and other sce…