collaborators

9 papers

cs.IR2026

CAMI: Cost-Aware Agent-Guided Multi-Indexing for Semantic Retrieval

Adnan Qidwai, Anand Eswaran, Sonam Mishra +2

RAG ingestion pipelines frequently augment search corpus index with semantic enrichment indices (e.g., synthetic queries or summaries generated from corpus chunks) that are subsequ…

cs.IR2026

Granite Embedding Multilingual R2 Models

Parul Awasthy, Aashka Trivedi, Yushu Yang +14

We introduce the multilingual Granite Embedding R2 models, a family of encoder-based embedding models for enterprise-scale dense retrieval across 200+ languages. Extending our Engl…

cs.IR2026

Influence Guided Sampling for Domain Adaptation of Text Retrievers

Meet Doshi, Vishwajeet Kumar, Yulong Li +1

General-purpose open-domain dense retrieval systems are usually trained with a large, eclectic mix of corpora and search tasks. How should these diverse corpora and tasks be sample…

cs.CL2026

LMK > CLS: Landmark Pooling for Dense Embeddings

Meet Doshi, Aashka Trivedi, Vishwajeet Kumar +5

Representation learning is central to many downstream tasks such as search, clustering, classification, and reranking. State-of-the-art sequence encoders typically collapse a varia…

cs.CL2025

Granite Embedding R2 Models

Parul Awasthy, Aashka Trivedi, Yulong Li +17

We introduce the Granite Embedding R2 models, a comprehensive family of high-performance English encoder-based embedding models engineered for enterprise-scale dense retrieval appl…

cs.IR2025

Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E5

Arkadeep Acharya, Rudra Murthy, Vishwajeet Kumar +1

Given the large number of Hindi speakers worldwide, there is a pressing need for robust and efficient information retrieval systems for Hindi. Despite ongoing research, comprehensi…