5 papers
Granite Embedding Multilingual R2 Models
Parul Awasthy, Aashka Trivedi, Yushu Yang +14
We introduce the multilingual Granite Embedding R2 models, a family of encoder-based embedding models for enterprise-scale dense retrieval across 200+ languages. Extending our Engl…
Influence Guided Sampling for Domain Adaptation of Text Retrievers
Meet Doshi, Vishwajeet Kumar, Yulong Li +1
General-purpose open-domain dense retrieval systems are usually trained with a large, eclectic mix of corpora and search tasks. How should these diverse corpora and tasks be sample…
LMK > CLS: Landmark Pooling for Dense Embeddings
Meet Doshi, Aashka Trivedi, Vishwajeet Kumar +5
Representation learning is central to many downstream tasks such as search, clustering, classification, and reranking. State-of-the-art sequence encoders typically collapse a varia…
Granite Embedding R2 Models
Parul Awasthy, Aashka Trivedi, Yulong Li +17
We introduce the Granite Embedding R2 models, a comprehensive family of high-performance English encoder-based embedding models engineered for enterprise-scale dense retrieval appl…
Pretraining Language Models Using Translationese
Meet Doshi, Raj Dabre, Pushpak Bhattacharyya
In this paper, we explore the utility of translationese as synthetic data created using machine translation for pre-training language models (LMs) for low-resource languages (LRLs)…