29 papers
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
Kaixin Ding, Xi Chen, Minghong Cai +9
Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllabi…
OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing
Chenxuan Miao, Yutong Feng, Yi Lu +6
Instruction-based video editing (IVE) is an emerging field with broad applications, yet evaluating editing models remains challenging. Existing benchmarks suffer from two major lim…
SURF: Signature-Retained Fast Video Generation
Kaixin Ding, Xi Chen, Sihui Ji +4
The demand for high-resolution video generation is growing rapidly. However, the generation resolution is severely constrained by slow inference speeds. For instance, Wan2.1 requir…
CoDance: An Unbind-Rebind Paradigm for Robust Multi-Subject Animation
Shuai Tan, Biao Gong, Ke Ma +5
Character image animation is gaining significant importance across various domains, driven by the demand for robust and flexible multi-subject rendering. While existing methods exc…
GDRO: Group-level Reward Post-training Suitable for Diffusion Models
Yiyang Wang, Xi Chen, Xiaogang Xu +2
Recent advancements adopt online reinforcement learning (RL) from LLMs to text-to-image rectified flow diffusion models for reward alignment. The use of group-level rewards success…
Alchemist: Unlocking Efficiency in Text-to-Image Model Training via Meta-Gradient Data Selection
Kaixin Ding, Yang Zhou, Xi Chen +5
Recent advances in Text-to-Image (T2I) generative models, such as Imagen, Stable Diffusion, and FLUX, have led to remarkable improvements in visual quality. However, their performa…