1 paper
Jiajun Chen, Sai Cheng, Yutao Yuan +4
Multimodal models integrating natural language and visual information have substantially improved generalization of representation models. However, their effectiveness significantl…