1 paper · 1 filter
Sen Li, Ruochen Wang, Cho-Jui Hsieh +2
Existing text-to-image models still struggle to generate images of multiple objects, especially in handling their spatial positions, relative sizes, overlapping, and attribute bind…