12 papers
Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation
Mining Tan, Yinuo Wang, Ziqi Zhou +6
Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances…
COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models
Ziqi Zhou, Weize Quan, Mining Tan +6
Composition is a high-level visual intent that governs where subjects are placed and how a scene is organized, yet current unified multimodal models remain unreliable at fine-grain…
Sissi: Zero-shot Style-guided Image Synthesis via Semantic-style Integration
Yingying Deng, Xiangyu He, Fan Tang +2
Text-guided image generation has advanced rapidly with large-scale diffusion models, yet achieving precise stylization with visual exemplars remains difficult. Existing approaches…
Inversion-Free Style Transfer with Dual Rectified Flows
Yingying Deng, Xiangyu He, Fan Tang +2
Style transfer, a pivotal task in image processing, synthesizes visually compelling images by seamlessly blending realistic content with artistic styles, enabling applications in p…
LumiSculpt: Enabling Consistent Portrait Lighting in Video Generation
Yuxin Zhang, Dandan Zheng, Biao Gong +5
Lighting plays a pivotal role in ensuring the naturalness and aesthetic quality of video generation. However, the impact of lighting is deeply coupled with other factors of videos,…
IP-Prompter: Training-Free Theme-Specific Image Generation via Dynamic Visual Prompting
Yuxin Zhang, Minyan Luo, Weiming Dong +6
The stories and characters that captivate us as we grow up shape unique fantasy worlds, with images serving as the primary medium for visually experiencing these realms. Personaliz…