1 paper
Jeeyung Kim, Erfan Esmaeili, Qiang Qiu
In text-to-image diffusion models, the cross-attention map of each text token indicates the specific image regions attended. Comparing these maps of syntactically related tokens pr…