3 citations · 3 across the 10 of their papers we have counts for
1 paper · 1 filter
Chungpa Lee, Jihoon Kwon, Kyle Min +1
Vision-language models map images and text into a joint embedding space. However, these embeddings often entangle multiple semantic features, which limits their interpretability an…