collaborators

5 papers

cs.IR2026

ColBERTSaR: Sparsified ColBERT Index via Product Quantization

Eugene Yang, Andrew Yates, Dawn Lawrie +3

While ColBERT is an effective neural retrieval architecture, it requires a heavy index structure to support candidate set retrieval based on approximated token embeddings, gatherin…

cs.IR2026

A Replicability Study of XTR

Rohan Jha, Reno Kriz, Benjamin Van Durme

The XTR (conteXtual Token Retrieval) algorithm is a modification to ColBERT retrieval that avoids the costly step of fully gathering and reranking the candidates' embeddings by imp…

cs.IR2026

A Brief Comparison of Training-Free Multi-Vector Sequence Compression Methods

Rohan Jha, Chunsheng Zuo, Reno Kriz +1

While multi-vector retrieval models outperform single-vector models of comparable size in retrieval quality, their practicality is limited by substantially larger index sizes, driv…

cs.IR2026

Multi-Vector Index Compression in Any Modality

Hanxiang Qin, Alexander Martin, Rohan Jha +3

We study efficient multi-vector retrieval for late interaction in any modality. Late interaction has emerged as a dominant paradigm for information retrieval in text, images, visua…

cs.CL2025

semantic-features: A User-Friendly Tool for Studying Contextual Word Embeddings in Interpretable Semantic Spaces

Jwalanthi Ranganathan, Rohan Jha, Kanishka Misra +1

We introduce semantic-features, an extensible, easy-to-use library based on Chronis et al. (2023) for studying contextualized word embeddings of LMs by projecting them into interpr…