1 paper · 1 filter
Manuel Tran, Yashin Dicente Cid, Amal Lahiani +3
Training multimodal foundation models is challenging due to the limited availability of multimodal datasets. While many public datasets pair images with text, few combine images wi…