3 papers · 1 filter
Spatially-Grounded Text-to-Video Generation via Inference-Time Gradient-Free Optimization
Guillaume Jeanneret, Mathis Koroglu, Hugo Caselles-Dupré +2
Diffusion Transformer Text-to-Video models have achieved remarkable synthesis quality, yet fine-grained spatial controllability remains a significant challenge. While existing trai…
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
Hugo Caselles-Dupré, Hugo Caselles-Dupré, Mathis Koroglu +3
Diffusion-based image-to-video (I2V) models are increasingly effective, yet they struggle to scale to ultra-high-resolution inputs (e.g., 4K). Generating videos at the model's nati…
OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models
Mathis Koroglu, Hugo Caselles-Dupré, Guillaume Jeanneret Sanmiguel +1
We consider the problem of text-to-video generation tasks with precise control for various applications such as camera movement control and video-to-video editing. Most methods tac…