activity
20242026
most citedZero-shot Image Editing with Reference Imitation

1 citations · 4 across the 29 of their papers we have counts for

collaborators

36 papers

cs.CV2026

OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing

Chenxuan Miao, Yutong Feng, Yi Lu +6

Instruction-based video editing (IVE) is an emerging field with broad applications, yet evaluating editing models remains challenging. Existing benchmarks suffer from two major lim…

cs.RO2026

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

Senqiao Yang, Chengyao Wang, Yuxin Chen +13

Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodie…

cs.CV2026

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

Kaixin Ding, Xi Chen, Minghong Cai +9

Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllabi…

cs.GR2026

SURF: Signature-Retained Fast Video Generation

Kaixin Ding, Xi Chen, Sihui Ji +4

The demand for high-resolution video generation is growing rapidly. However, the generation resolution is severely constrained by slow inference speeds. For instance, Wan2.1 requir…

cs.CV2026

CoDance: An Unbind-Rebind Paradigm for Robust Multi-Subject Animation

Shuai Tan, Biao Gong, Ke Ma +5

Character image animation is gaining significant importance across various domains, driven by the demand for robust and flexible multi-subject rendering. While existing methods exc…

cs.LG2026

GDRO: Group-level Reward Post-training Suitable for Diffusion Models

Yiyang Wang, Xi Chen, Xiaogang Xu +2

Recent advancements adopt online reinforcement learning (RL) from LLMs to text-to-image rectified flow diffusion models for reward alignment. The use of group-level rewards success…