1 paper
Jiawei Ma, Yulei Niu, Shiyuan Huang +2
Language has been useful in extending the vision encoder to data from diverse distributions without empirical discovery in training domains. However, as the image description is mo…