1 paper · 1 filter
Lexiang Hu, Youze Xue, Dian Li +2
Multimodal embeddings serve as a bridge for aligning vision and language, with the two primary implementations -- CLIP-based and MLLM-based embedding models -- both limited to capt…