collaborators

9 papers

cs.CV2026

UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

Liming Tan, Ye Chen, Hao Zhang +3

Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occl…

cs.CV2026

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration

Ye Chen, Xuanhong Chen, Yupeng Zhu +23

The paper proposes the World Narrative Model, a framework that separates the specification of a 4D physical scene (geometry, motion, camera, lighting) from pixel generation, enabli…

cs.CV2026

Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information Flows

Xiang Yang, Feifei Li, Mi Zhang +4

Diffusion transformers (DiTs) equipped with multimodal attention (MM-Attn) have become a dominant paradigm for image generation. However, preventing the generation of harmful conte…

cs.CV2026

Broken Memories: Detecting and Mitigating Memorization in Diffusion Models with Degraded Generations

Yuanmin Huang, Mi Zhang, Chen Chen +4

While diffusion models excel at generating high-quality images, their tendency to memorize training data poses significant privacy and copyright risks. In this work, we for the fir…

cs.CV2026

VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation

Mengtian Li, Yuwei Lu, Feifei Li +3

Cinematic camera control relies on a tight feedback loop between director and cinematographer, where camera motion and framing are continuously reviewed and refined. Recent generat…

cs.CV2026

SafeRoPE: Risk-specific Head-wise Embedding Rotation for Safe Generation in Rectified Flow Transformers

Xiang Yang, Feifei Li, Mi Zhang +3

Recent Text-to-Image (T2I) models based on rectified-flow transformers (e.g., SD3, FLUX) achieve high generative fidelity but remain vulnerable to unsafe semantics, especially when…