1 paper
Atharvan Dogra, Deeksha Varshney, Ashwin Kalyan +2
The generation of effective latent representations and their subsequent refinement to incorporate precise information is an essential prerequisite for Vision-Language Understanding…