collaborators

6 papers

cs.CV2025

CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching

Chen Chen, Pengsheng Guo, Liangchen Song +7

Conditional generative modeling aims to learn a conditional data distribution from samples containing data-condition pairs. For this, diffusion and flow-based methods have attained…

cs.CV2025

Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing

Yusu Qian, Eli Bocek-Rivele, Liangchen Song +5

Recent advances in multimodal models have demonstrated remarkable text-guided image editing capabilities, with systems like GPT-4o and Nano-Banana setting new benchmarks. However,…

cs.CV2025

Autoregressive Video Generation beyond Next Frames Prediction

Sucheng Ren, Chen Chen, Zhenbang Wang +5

Autoregressive models for video generation typically operate frame-by-frame, extending next-token prediction from language to video's temporal dimension. We question that unlike wo…

cs.CV2025

AToken: A Unified Tokenizer for Vision

Jiasen Lu, Liangchen Song, Mingze Xu +5

We present AToken, the first unified visual tokenizer that achieves both high-fidelity reconstruction and semantic understanding across images, videos, and 3D assets. Unlike existi…

cs.CV2025

Score Distillation of Flow Matching Models

Mingyuan Zhou, Yi Gu, Huangjie Zheng +5

Diffusion models achieve high-quality image generation but are limited by slow iterative sampling. Distillation methods alleviate this by enabling one- or few-step generation. Flow…

cs.CV2024

STIV: Scalable Text and Image Conditioned Video Generation

Zongyu Lin, Wei Liu, Chen Chen +13

The field of video generation has made remarkable advancements, yet there remains a pressing need for a clear, systematic recipe that can guide the development of robust and scalab…