1 paper
Dahee Kwon, Haeun Lee, Jaesik Choi
Recent text-to-image models built on large-scale Transformer backbones and flow-based objectives deliver strong text-image alignment and high visual quality, yet often produce over…