9 papers
UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation
Liming Tan, Ye Chen, Hao Zhang +3
Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occl…
World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration
Ye Chen, Xuanhong Chen, Yupeng Zhu +23
The paper proposes the World Narrative Model, a framework that separates the specification of a 4D physical scene (geometry, motion, camera, lighting) from pixel generation, enabli…
Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information Flows
Xiang Yang, Feifei Li, Mi Zhang +4
Diffusion transformers (DiTs) equipped with multimodal attention (MM-Attn) have become a dominant paradigm for image generation. However, preventing the generation of harmful conte…
Broken Memories: Detecting and Mitigating Memorization in Diffusion Models with Degraded Generations
Yuanmin Huang, Mi Zhang, Chen Chen +4
While diffusion models excel at generating high-quality images, their tendency to memorize training data poses significant privacy and copyright risks. In this work, we for the fir…
VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation
Mengtian Li, Yuwei Lu, Feifei Li +3
Cinematic camera control relies on a tight feedback loop between director and cinematographer, where camera motion and framing are continuously reviewed and refined. Recent generat…
SafeRoPE: Risk-specific Head-wise Embedding Rotation for Safe Generation in Rectified Flow Transformers
Xiang Yang, Feifei Li, Mi Zhang +3
Recent Text-to-Image (T2I) models based on rectified-flow transformers (e.g., SD3, FLUX) achieve high generative fidelity but remain vulnerable to unsafe semantics, especially when…