1 paper · 1 filter
Kin Ian Lo
Dual-encoder models such as CLIP score an image-caption pair by a single inner product of two independently computed unit vectors, and fail at binding, often scoring near chance wh…