8 papers
Granite Embedding Multilingual R2 Models
Parul Awasthy, Aashka Trivedi, Yushu Yang +14
We introduce the multilingual Granite Embedding R2 models, a family of encoder-based embedding models for enterprise-scale dense retrieval across 200+ languages. Extending our Engl…
Influence Guided Sampling for Domain Adaptation of Text Retrievers
Meet Doshi, Vishwajeet Kumar, Yulong Li +1
General-purpose open-domain dense retrieval systems are usually trained with a large, eclectic mix of corpora and search tasks. How should these diverse corpora and tasks be sample…
LMK > CLS: Landmark Pooling for Dense Embeddings
Meet Doshi, Aashka Trivedi, Vishwajeet Kumar +5
Representation learning is central to many downstream tasks such as search, clustering, classification, and reranking. State-of-the-art sequence encoders typically collapse a varia…
Granite Embedding R2 Models
Parul Awasthy, Aashka Trivedi, Yulong Li +17
We introduce the Granite Embedding R2 models, a comprehensive family of high-performance English encoder-based embedding models engineered for enterprise-scale dense retrieval appl…
Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E5
Arkadeep Acharya, Rudra Murthy, Vishwajeet Kumar +1
Given the large number of Hindi speakers worldwide, there is a pressing need for robust and efficient information retrieval systems for Hindi. Despite ongoing research, comprehensi…
Granite Embedding Models
Parul Awasthy, Aashka Trivedi, Yulong Li +19
We introduce the Granite Embedding models, a family of encoder-based embedding models designed for retrieval tasks, spanning dense-retrieval and sparse retrieval architectures, wit…