1 paper
Anthony Meng Huat Tiong, Junqi Zhao, Boyang Li +3
Vision-language (VL) models, pretrained on colossal image-text datasets, have attained broad VL competence that is difficult to evaluate. A common belief is that a small number of…