6 papers
CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities
Yue Gong, Shanyuan Liu, Liuzhuozheng Li +7
We proposed the Chinese Text Adapter-Flux (CTA-Flux). An adaptation method fits the Chinese text inputs to Flux, a powerful text-to-image (TTI) generative model initially trained o…
NanoControl: A Lightweight Framework for Precise and Efficient Control in Diffusion Transformer
Shanyuan Liu, Jian Zhu, Junda Lu +8
Diffusion Transformers (DiTs) have demonstrated exceptional capabilities in text-to-image synthesis. However, in the domain of controllable text-to-image generation using DiTs, mos…
FLUX-Makeup: High-Fidelity, Identity-Consistent, and Robust Makeup Transfer via Diffusion Transformer
Jian Zhu, Shanyuan Liu, Liuzhuozheng Li +9
Makeup transfer aims to apply the makeup style from a reference face to a target face and has been increasingly adopted in practical applications. Existing GAN-based approaches typ…
PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language Models
Runze He, Bo Cheng, Yuhang Ma +7
In this paper, we propose a unified layout planning and image generation model, PlanGen, which can pre-plan spatial layout conditions before generating images. Unlike previous diff…
U-StyDiT: Ultra-high Quality Artistic Style Transfer Using Diffusion Transformers
Zhanjie Zhang, Ao Ma, Ke Cao +6
Ultra-high quality artistic style transfer refers to repainting an ultra-high quality content image using the style information learned from the style image. Existing artistic styl…
WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation
Jing Wang, Ao Ma, Ke Cao +9
Recent rapid advancements in text-to-video (T2V) generation, such as SoRA and Kling, have shown great potential for building world simulators. However, current T2V models struggle…