1 paper
Ahmad Elallaf, Yu Zhang, Yuktha Priya Masupalli +4
Vision-language foundation models have emerged as powerful general-purpose representation learners with strong potential for multimodal understanding, but their deterministic embed…