5 papers
PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation
Yanjie Pan, Qingdong He, Zhengkai Jiang +9
Recent advances in diffusion-based text-to-image generation have demonstrated promising results through visual condition control. However, existing ControlNet-like methods struggle…
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
Shuang Chen, Quanxin Shou, Hangting Chen +16
Unified multimodal models provide a natural and promising architecture for understanding diverse and complex real-world knowledge while generating high-quality images. However, the…
Investigating Text Insulation and Attention Mechanisms for Complex Visual Text Generation
Ying Tai, Nikai Du, Rui Xie +5
In this paper, we present TextCrafter, a Complex Visual Text Generation (CVTG) framework inspired by selective visual attention in cognitive science, and introduce the "Text Insula…
RealVVT: Towards Photorealistic Video Virtual Try-on via Spatio-Temporal Consistency
Siqi Li, Zhengkai Jiang, Jiawei Zhou +3
Virtual try-on has emerged as a pivotal task at the intersection of computer vision and fashion, aimed at digitally simulating how clothing items fit on the human body. Despite not…
Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement
Zhennan Chen, Yajie Li, Haofan Wang +6
Regional prompting, or compositional generation, which enables fine-grained spatial control, has gained increasing attention for its practicality in real-world applications. Howeve…