3 citations · 5 across the 11 of their papers we have counts for
1 paper · 1 filter
Di Wu, Yixin Wan, Kai-Wei Chang
Text-to-image retrieval (T2I retrieval) remains challenging because cross-modal embeddings often behave as bags of concepts, underrepresenting structured visual relationships such…