4 papers
JoLT: Joint Latent Trajectories for Context-Guided High-Resolution Tiled Generation
Mathis Koroglu, Guillaume Jeanneret, Hugo Caselles-Dupré +2
Although text-to-image generative models produce impressive results, they struggle to generate densely detailed, high-resolution (HR) images. Current literature addresses this issu…
Spatially-Grounded Text-to-Video Generation via Inference-Time Gradient-Free Optimization
Guillaume Jeanneret, Mathis Koroglu, Hugo Caselles-Dupré +2
Diffusion Transformer Text-to-Video models have achieved remarkable synthesis quality, yet fine-grained spatial controllability remains a significant challenge. While existing trai…
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
Hugo Caselles-Dupré, Mathis Koroglu, Guillaume Jeanneret +2
Diffusion-based image-to-video (I2V) models are increasingly effective, yet they struggle to scale to ultra-high-resolution inputs (e.g., 4K). Generating videos at the model's nati…
OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models
Mathis Koroglu, Hugo Caselles-Dupré, Guillaume Jeanneret Sanmiguel +1
We consider the problem of text-to-video generation tasks with precise control for various applications such as camera movement control and video-to-video editing. Most methods tac…