1 paper · 1 filter
Weili Nie, Sifei Liu, Morteza Mardani +3
Existing text-to-image models struggle to follow complex text prompts, raising the need for extra grounding inputs for better controllability. In this work, we propose to decompose…