1 paper
N. K. B. M. P. K. B. Narasinghe, Uthayasanker Thayasivam
Large-scale multimodal foundation models, particularly Contrastive Captioners (CoCa), have achieved state-of-the-art results by unifying contrastive alignment with generative capti…