1 paper
Youngtaek Oh, Jae Won Cho, Dong-Jin Kim +2
In this paper, we propose a new method to enhance compositional understanding in pre-trained vision and language models (VLMs) without sacrificing performance in zero-shot multi-mo…