21 citations · 22 across the 5 of their papers we have counts for
1 paper · 1 filter
Souhail Hadgi, Luca Moschella, Andrea Santilli +5
Recent works have shown that, when trained at scale, uni-modal 2D vision and text encoders converge to learned features that share remarkable structural properties, despite arising…