1 citations · 1 across the 5 of their papers we have counts for
6 papers
Are Image-to-Video Models Good Zero-Shot Image Editors?
Zechuan Zhang, Zhenyuan Chen, Zongxin Yang +1
Large-scale video diffusion models show strong world simulation and temporal reasoning abilities, but their use as zero-shot image editors remains underexplored. We introduce IF-Ed…
BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
Dewei Zhou, Mingwei Li, Zongxin Yang +5
Conditional image generation enhances text-to-image synthesis with structural, spatial, or stylistic priors, but current methods face challenges in handling conflicts between sourc…
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
Dewei Zhou, Mingwei Li, Zongxin Yang +1
Image-conditioned generation methods, such as depth- and canny-conditioned approaches, have demonstrated remarkable abilities for precise image synthesis. However, existing models…
In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
Zechuan Zhang, Ji Xie, Yu Lu +2
Instruction-based image editing enables precise modifications via natural language prompts, but existing methods face a precision-efficiency tradeoff: fine-tuning demands massive d…
3D Object Manipulation in a Single Image using Generative Models
Ruisi Zhao, Zechuan Zhang, Zongxin Yang +1
Object manipulation in images aims to not only edit the object's presentation but also gift objects with motion. Previous methods encountered challenges in concurrently handling st…
3DIS-FLUX: simple and efficient multi-instance generation with DiT rendering
Dewei Zhou, Ji Xie, Zongxin Yang +1
The growing demand for controllable outputs in text-to-image generation has driven significant advancements in multi-instance generation (MIG), enabling users to define both instan…