papers

Publications (15)

cs.IR2025

Granite Embedding Models

Parul Awasthy, Aashka Trivedi, Yulong Li +19

We introduce the Granite Embedding models, a family of encoder-based embedding models designed for retrieval tasks, spanning dense-retrieval and sparse retrieval architectures, wit…

cs.CL2025

MILU: A Multi-task Indic Language Understanding Benchmark

Sshubam Verma, Mohammed Safi Ur Rahman Khan, Vishwajeet Kumar +2

Evaluating Large Language Models (LLMs) in low-resource and linguistically diverse languages remains a significant challenge in NLP, particularly for languages using non-Latin scri…

cs.LG2025

INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages

Abhishek Kumar Singh, Vishwajeet kumar, Rudra Murthy +3

Large Language Models (LLMs) perform well on unseen tasks in English, but their abilities in non English languages are less explored due to limited benchmarks and training data. To…

cs.CL2023

PrimeQA: The Prime Repository for State-of-the-Art Multilingual Question Answering Research and Development

Avirup Sil, Jaydeep Sen, Bhavani Iyer +12

The field of Question Answering (QA) has made remarkable progress in recent years, thanks to the advent of large pre-trained language models, newer realistic benchmark datasets wit…

cs.IR2026

Influence Guided Sampling for Domain Adaptation of Text Retrievers

Meet Doshi, Vishwajeet Kumar, Yulong Li +1

General-purpose open-domain dense retrieval systems are usually trained with a large, eclectic mix of corpora and search tasks. How should these diverse corpora and tasks be sample…

cs.IR2026

CAMI: Cost-Aware Agent-Guided Multi-Indexing for Semantic Retrieval

Adnan Qidwai, Anand Eswaran, Sonam Mishra +2

RAG ingestion pipelines frequently augment search corpus index with semantic enrichment indices (e.g., synthetic queries or summaries generated from corpus chunks) that are subsequ…

cs.IR2026

Granite Embedding Multilingual R2 Models

Parul Awasthy, Aashka Trivedi, Yushu Yang +14

We introduce the multilingual Granite Embedding R2 models, a family of encoder-based embedding models for enterprise-scale dense retrieval across 200+ languages. Extending our Engl…

cs.IR2024

Hindi-BEIR : A Large Scale Retrieval Benchmark in Hindi

Arkadeep Acharya, Rudra Murthy, Vishwajeet Kumar +1

Given the large number of Hindi speakers worldwide, there is a pressing need for robust and efficient information retrieval systems for Hindi. Despite ongoing research, there is a…

cs.CL2021

Topic Transferable Table Question Answering

Saneem Ahmed Chemmengath, Vishwajeet Kumar, Samarth Bharadwaj +5

Weakly-supervised table question-answering(TableQA) models have achieved state-of-art performance by using pre-trained BERT transformer to jointly encoding a question and a table t…

cs.IR2024

Mistral-SPLADE: LLMs for better Learned Sparse Retrieval

Meet Doshi, Vishwajeet Kumar, Rudra Murthy +2

Learned Sparse Retrievers (LSR) have evolved into an effective retrieval strategy that can bridge the gap between traditional keyword-based sparse retrievers and embedding-based de…

cs.CL2021

AIT-QA: Question Answering Dataset over Complex Tables in the Airline Industry

Yannis Katsis, Saneem Chemmengath, Vishwajeet Kumar +8

Recent advances in transformers have enabled Table Question Answering (Table QA) systems to achieve high accuracy and SOTA results on open domain datasets like WikiTableQuestions a…

cs.CL2026

LMK > CLS: Landmark Pooling for Dense Embeddings

Meet Doshi, Aashka Trivedi, Vishwajeet Kumar +5

Representation learning is central to many downstream tasks such as search, clustering, classification, and reranking. State-of-the-art sequence encoders typically collapse a varia…

cs.CL2025

Granite Embedding R2 Models

Parul Awasthy, Aashka Trivedi, Yulong Li +17

We introduce the Granite Embedding R2 models, a comprehensive family of high-performance English encoder-based embedding models engineered for enterprise-scale dense retrieval appl…

cs.IR2025

Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E5

Arkadeep Acharya, Rudra Murthy, Vishwajeet Kumar +1

Given the large number of Hindi speakers worldwide, there is a pressing need for robust and efficient information retrieval systems for Hindi. Despite ongoing research, comprehensi…

cs.CL2023

Multi-Row, Multi-Span Distant Supervision For Table+Text Question

Vishwajeet Kumar, Yash Gupta, Saneem Chemmengath +4

Question answering (QA) over tables and linked text, also called TextTableQA, has witnessed significant research in recent years, as tables are often found embedded in documents al…