1 paper
Daiki Yoshikawa, Takashi Matsubara
Vision-language models have achieved remarkable success in multi-modal representation learning from large-scale pairs of visual scenes and linguistic descriptions. However, they st…