1 paper · 1 filter
Shiyu Li, Zhiyuan Hu, Yifan Wang +3
Omni-modal retrieval promises a single embedding space for text, image, video, document, and audio inputs, but building such a unified retriever is difficult since these modalities…