1 paper
Akayou A. Kitessa, Yijun Zhao
Vision-language models map visual features into a shared embedding space through learned projection layers, yet it remains unclear how these transformations alter the structure of…