4 papers
ContextDrag: Precise Drag-Based Image Editing via Context-Preserving Token Injection and Position-Aligned Attention
Huiguo He, Pengyu Yan, Ziqi Yi +6
Drag-based image editing enables intuitive visual manipulation through point-based drag operations. Existing methods mainly rely on diffusion inversion or pixel-space warping with…
SceneExpander: Text-Guided 3D Scene Expansion via Free-Form View Insertion
Zijian He, Renjie Liu, Yihao Wang +5
World building with 3D scene representations is increasingly important for content creation, simulation, and interactive experiences, yet real workflows are inherently iterative: c…
Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation Adapter
Weizhi Zhong, Huan Yang, Zheng Liu +5
Personalized text-to-image generation aims to synthesize images of user-provided concepts in diverse contexts. Despite recent progress in multi-concept personalization, most are li…
Style-Preserving Lip Sync via Audio-Aware Style Reference
Weizhi Zhong, Jichang Li, Yinqi Cai +4
Audio-driven lip sync has recently drawn significant attention due to its widespread application in the multimedia domain. Individuals exhibit distinct lip shapes when speaking the…