1 paper · 1 filter
Jiabao Shi, Minfeng Qi, Lefeng Zhang +5
Multimodal text-to-image generation remains constrained by the difficulty of maintaining semantic alignment and professional-level detail across diverse visual domains. We propose…