7 papers
Douyin Multimodal Embedding Model Technical Report
Haonan Chen, Chu Li, Zhicheng Wang +4
Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and recommendation and underp…
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
Haonan Chen, Sicheng Gao, Radu Timofte +2
Modern information systems often involve different types of items, e.g., a text query, an image, a video clip, or an audio segment. This motivates omni-modal embedding models that…
Chain-of-Retrieval Augmented Generation
Liang Wang, Haonan Chen, Nan Yang +3
This paper introduces an approach for training o1-like RAG models that retrieve and reason over relevant information step by step before generating the final answer. Conventional R…
A Survey of Conversational Search
Fengran Mo, Kelong Mao, Ziliang Zhao +7
As a cornerstone of modern information access, search engines have become indispensable in everyday life. With the rapid advancements in AI and natural language processing (NLP) te…
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
Haonan Chen, Hong Liu, Yuping Luo +4
Multimodal embedding models, built upon causal Vision Language Models (VLMs), have shown promise in various tasks. However, current approaches face three key limitations: the use o…
mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data
Haonan Chen, Liang Wang, Nan Yang +4
Multimodal embedding models have gained significant attention for their ability to map data from different modalities, such as text and images, into a unified representation space.…