2 citations · 4 across the 12 of their papers we have counts for
12 papers
Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
Jing Tan, Zhaoyang Zhang, Yantao Shen +6
We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text-instructed spatial transformation of objects within scenes. Spatially manipulating objects…
IC-Custom: Diverse Image Customization via In-Context Learning
Yaowei Li, Xiaoyu Li, Zhaoyang Zhang +11
Image customization, a crucial technique for industrial media production, aims to generate content that is consistent with reference images. However, current approaches conventiona…
FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios
Shiyi Zhang, Junhao Zhuang, Zhaoyang Zhang +2
Action customization involves generating videos where the subject performs actions dictated by input control signals. Current methods use pose-guided or global motion customization…
Cobra: Efficient Line Art COlorization with BRoAder References
Junhao Zhuang, Lingen Li, Xuan Ju +3
The comic production industry requires reference-based line art colorization with high accuracy, efficiency, contextual consistency, and flexible control. A comic page often involv…
BlobCtrl: Taming Controllable Blob for Element-level Image Editing
Yaowei Li, Lingen Li, Zhaoyang Zhang +6
As user expectations for image editing continue to rise, the demand for flexible, fine-grained manipulation of specific visual elements presents a challenge for current diffusion-b…
Consistent Human Image and Video Generation with Spatially Conditioned Diffusion
Mingdeng Cao, Chong Mou, Ziyang Yuan +4
Consistent human-centric image and video synthesis aims to generate images or videos with new poses while preserving appearance consistency with a given reference image, which is c…