3 papers
cs.CV2026
MosaicMem: Hybrid Spatial Memory for Controllable Video World Models
Wei Yu, Runjia Qian, Yumeng Li +8
Video diffusion models are moving beyond short, plausible clips toward world simulators that must remain consistent under camera motion, revisits, and intervention. Yet spatial mem…
cs.CV2026
PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic Trajectories
Gemma Canet Tarrés, Manel Baradad, Francesc Moreno-Noguer +1
Recent advances in generative AI have dramatically improved photorealistic image synthesis, yet they fall short for studio-level multi-object compositing. This task demands simulta…
cs.CV2025
VSTAR: Generative Temporal Nursing for Longer Dynamic Video Synthesis
Yumeng Li, William Beluch, Margret Keuper +2
Despite tremendous progress in the field of text-to-video (T2V) synthesis, open-sourced T2V diffusion models struggle to generate longer videos with dynamically varying and evolvin…