1 paper
Tsung-Wei Ke, Sangwoo Mo, Stella X. Yu
Large vision and language models learned directly through image-text associations often lack detailed visual substantiation, whereas image segmentation tasks are treated separately…