1 paper
Jingxiang Sun, Chao Liao, Zhengxiong Luo +6
Current text-to-image systems typically employ a "text encoder plus diffusion decoder" paradigm, in which text semantics directly modulate continuous latent noise. Despite their su…