works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

6 papers

cs.LG2026

Amortized Moment Matching for Visual Generation

Wenze Liu, Xintao Wang, Pengfei Wan +1

The paper introduces amortized moment matching, using neural networks to learn data moments as training signals, and proposes the Amortized Fréchet Distance loss to improve one-ste…

cs.CV2026

VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning

Minghong Cai, Qiulin Wang, Zongli Ye +7

Existing controllable video generation methods are typically designed for rigid, task-specific settings, such as first-frame image-to-video, inpainting, or interpolation, treating…

cs.CV2025

In-Context Audio Control of Video Diffusion Transformers

Wenze Liu, Weicai Ye, Minghong Cai +3

Recent advancements in video generation have seen a shift towards unified, transformer-based foundation models that can handle multiple conditional inputs in-context. However, thes…

cs.LG2025

Learning to Integrate Diffusion ODEs by Averaging the Derivatives

Wenze Liu, Xiangyu Yue

To accelerate diffusion model inference, numerical solvers perform poorly at extremely small steps, while distillation techniques often introduce complexity and instability. This w…

cs.CV2025

DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation

Minghong Cai, Xiaodong Cun, Xiaoyu Li +5

Sora-like video generation models have achieved remarkable progress with a Multi-Modal Diffusion Transformer MM-DiT architecture. However, the current video generation models predo…

cs.CV2024

Customize Your Visual Autoregressive Recipe with Set Autoregressive Modeling

Wenze Liu, Le Zhuo, Yi Xin +3

We introduce a new paradigm for AutoRegressive (AR) image generation, termed Set AutoRegressive Modeling (SAR). SAR generalizes the conventional AR to the next-set setting, i.e., s…