activity
20242026
most citedMiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions

2 citations · 4 across the 12 of their papers we have counts for

collaborators

12 papers

cs.CV2026

Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes

Jing Tan, Zhaoyang Zhang, Yantao Shen +6

We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text-instructed spatial transformation of objects within scenes. Spatially manipulating objects…

cs.CV2025

IC-Custom: Diverse Image Customization via In-Context Learning

Yaowei Li, Xiaoyu Li, Zhaoyang Zhang +11

Image customization, a crucial technique for industrial media production, aims to generate content that is consistent with reference images. However, current approaches conventiona…

cs.CV2025

FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios

Shiyi Zhang, Junhao Zhuang, Zhaoyang Zhang +2

Action customization involves generating videos where the subject performs actions dictated by input control signals. Current methods use pose-guided or global motion customization…

cs.CV2025

Cobra: Efficient Line Art COlorization with BRoAder References

Junhao Zhuang, Lingen Li, Xuan Ju +3

The comic production industry requires reference-based line art colorization with high accuracy, efficiency, contextual consistency, and flexible control. A comic page often involv…

cs.CV2025

BlobCtrl: Taming Controllable Blob for Element-level Image Editing

Yaowei Li, Lingen Li, Zhaoyang Zhang +6

As user expectations for image editing continue to rise, the demand for flexible, fine-grained manipulation of specific visual elements presents a challenge for current diffusion-b…

cs.CV2024

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion

Mingdeng Cao, Chong Mou, Ziyang Yuan +4

Consistent human-centric image and video synthesis aims to generate images or videos with new poses while preserving appearance consistency with a given reference image, which is c…