2 papers
cs.CV2026
Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation
Yushe Cao, Dianxi Shi, Xing Fu +5
While significant progress has been achieved in multimodal facial generation using semantic masks and textual descriptions, conventional feature fusion approaches often fail to ena…
cs.CV2025
Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
Xuechao Zou, Shun Zhang, Xing Fu +6
Controllable face generation poses critical challenges in generative modeling due to the intricate balance required between semantic controllability and photorealism. While existin…