6 papers
Adversarial Flow Models
Shanchuan Lin, Ceyuan Yang, Zhijie Lin +2
We present adversarial flow models, a class of generative models that belongs to both the adversarial and flow families. Our method supports native one-step and multi-step generati…
Morphology-Consistent Humanoid Interaction through Robot-Centric Video Synthesis
Weisheng Xu, Jian Li, Yi Gu +12
Equipping humanoid robots with versatile interaction skills typically requires either extensive policy training or explicit human-to-robot motion retargeting. However, learning-bas…
Accelerating Diffusion Decoders via Multi-Scale Sampling and One-Step Distillation
Chuhan Wang, Hao Chen
Image tokenization plays a central role in modern generative modeling by mapping visual inputs into compact representations that serve as an intermediate signal between pixels and…
SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training
Jianyi Wang, Shanchuan Lin, Zhijie Lin +10
Recent advances in diffusion-based video restoration (VR) demonstrate significant improvement in visual quality, yet yield a prohibitive computational cost during inference. While…
SkipSR: Faster Super Resolution with Token Skipping
Rohan Choudhury, Shanchuan Lin, Jianyi Wang +6
Diffusion-based super-resolution (SR) is a key component in video generation and video restoration, but is slow and expensive, limiting scalability to higher resolutions and longer…
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
Jiaming Han, Hao Chen, Yang Zhao +6
This paper presents a multimodal framework that attempts to unify visual understanding and generation within a shared discrete semantic representation. At its core is the Text-Alig…