1 paper
Baiqi Li, Zhiqiu Lin, Deepak Pathak +8
While text-to-visual models now produce photo-realistic images and videos, they struggle with compositional text prompts involving attributes, relationships, and higher-order reaso…