5 citations · 5 across the 1 of their papers we have counts for
1 paper
Chuofan Ma, Yi Jiang, Xin Wen +2
Deriving reliable region-word alignment from image-text pairs is critical to learn object-level vision-language representations for open-vocabulary object detection. Existing metho…