17 papers
UENR-600K: A Large-Scale Physically Grounded Dataset for Nighttime Video Deraining
Pei Yang, Hai Ci, Beibei Lin +2
Nighttime video deraining is uniquely challenging because raindrops interact with artificial lighting. Unlike daytime white rain, nighttime rain takes on various colors and appears…
Loom: Diffusion-Transformer for Interleaved Generation
Mingcheng Ye, Jiaming Liu, Yiren Song
Interleaved text-image generation aims to jointly produce coherent visual frames and aligned textual descriptions within a single sequence, enabling tasks such as style transfer, c…
X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
Pei Yang, Hai Ci, Yiren Song +1
The advancement of embodied AI has unlocked significant potential for intelligent humanoid robots. However, progress in both Vision-Language-Action (VLA) models and world models is…
TokenPure: Watermark Removal through Tokenized Appearance and Structural Guidance
Pei Yang, Yepeng Liu, Kelly Peng +2
In the digital economy era, digital watermarking serves as a critical basis for ownership proof of massive replicable content, including AI-generated and other virtual assets. Desi…
WordCon: Word-level Typography Control in Scene Text Rendering
Wenda Shi, Yiren Song, Zihan Rao +3
Achieving precise word-level typography control within generated images remains a persistent challenge. To address it, we newly construct a word-level controlled scene text dataset…
OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data
Yiren Song, Cheng Liu, Mike Zheng Shou
Diffusion models have advanced image stylization significantly, yet two core challenges persist: (1) maintaining consistent stylization in complex scenes, particularly identity, co…