15 citations · 18 across the 9 of their papers we have counts for
12 papers · 1 filter
MultiAnimate: A Unified Framework for Controllable Multi-Character Animation
Zhongyi Zhang, Guangyuan Wang, Li Hu +6
Recent advances in generative models and technological innovations have significantly addressed the fundamental challenges of character image animation. However, existing approache…
Rethinking Multi-Condition DiTs: Eliminating Redundant Attention via Position-Alignment and Keyword-Scoping
Chao Zhou, Tianyi Wei, Yiling Chen +2
While modern text-to-image models excel at prompt-based generation, they often lack the fine-grained control necessary for specific user requirements like spatial layouts or subjec…
Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
Chao Zhou, Tianyi Wei, Nenghai Yu
Recent advancements in unified image generation models, such as OmniGen, have enabled the handling of diverse image generation and editing tasks within a single framework, acceptin…
WordCraft: Interactive Artistic Typography with Attention Awareness and Noise Blending
Zhe Wang, Jingbo Zhang, Tianyi Wei +2
Artistic typography aims to stylize input characters with visual effects that are both creative and legible. Traditional approaches rely heavily on manual design, while recent gene…
FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing
Tianyi Wei, Yifan Zhou, Dongdong Chen +1
The integration of Rotary Position Embedding (RoPE) in Multimodal Diffusion Transformer (MMDiT) has significantly enhanced text-to-image generation quality. However, the fundamenta…
Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation
Tianyi Wei, Dongdong Chen, Yifan Zhou +1
Representing the cutting-edge technique of text-to-image models, the latest Multimodal Diffusion Transformer (MMDiT) largely mitigates many generation issues existing in previous m…