1 paper · 1 filter
Souhail Hadgi, Luca Moschella, Andrea Santilli +5
Recent works have shown that, when trained at scale, uni-modal 2D vision and text encoders converge to learned features that share remarkable structural properties, despite arising…