1 paper
Z. Zhang, B. Liu, J. Bao +3
Recent text-to-image generation favors various forms of spatial conditions, e.g., masks, bounding boxes, and key points. However, the majority of the prior art requires form-specif…