1 paper · 1 filter
Xiaorui Ma, Haoran Xie, S. Joe Qin
The integration of vision-language modalities has been a significant focus in multimodal learning, traditionally relying on Vision-Language Pretrained Models. However, with the adv…