7 papers
Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation
Minyan Luo, Yuxin Zhang, Yifei Li +5
Narrative-driven product photography has become a prevalent paradigm in modern marketing, as coherent visual storytelling helps convey product value and establishes emotional engag…
SongEcho: Towards Cover Song Generation via Instance-Adaptive Element-wise Linear Modulation
Sifei Li, Yang Li, Zizhou Wang +5
Cover songs constitute a vital aspect of musical culture, preserving the core melody of an original composition while reinterpreting it to infuse novel emotional depth and thematic…
Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning
Yu Xu, Yuxin Zhang, Juan Cao +5
A visual metaphor constitutes a high-order form of human creativity, employing cross-domain semantic fusion to transform abstract concepts into impactful visual rhetoric. Despite t…
In-Context Brush: Zero-shot Customized Subject Insertion with Context-Aware Latent Space Manipulation
Yu Xu, Fan Tang, You Wu +6
Recent advances in diffusion models have enhanced multimodal-guided visual generation, enabling customized subject insertion that seamlessly "brushes" user-specified objects into a…
IP-Prompter: Training-Free Theme-Specific Image Generation via Dynamic Visual Prompting
Yuxin Zhang, Minyan Luo, Weiming Dong +6
The stories and characters that captivate us as we grow up shape unique fantasy worlds, with images serving as the primary medium for visually experiencing these realms. Personaliz…
HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads
Yu Xu, Fan Tang, Juan Cao +5
Diffusion Transformers (DiTs) have exhibited robust capabilities in image generation tasks. However, accurate text-guided image editing for multimodal DiTs (MM-DiTs) still poses a…