collaborators

10 papers

cs.IR2026

Overview of the TREC 2025 Retrieval Augmented Generation (RAG) Track

Shivani Upadhyay, Nandan Thakur, Ronak Pradeep +3

The second edition of the TREC Retrieval Augmented Generation (RAG) Track advances research on systems that integrate retrieval and generation to address complex, real-world inform…

cs.CL2025

MMTEB: Massive Multilingual Text Embedding Benchmark

Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83

Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…

cs.IR2025

On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools

Shivani Upadhyay, Messiah Ataey, Syed Shariyar Murtaza +2

The proliferation of complex structured data in hybrid sources, such as PDF documents and web pages, presents unique challenges for current Large Language Models (LLMs) and Multi-m…

cs.IR2025

Overview of the TREC 2023 deep learning track

Nick Craswell, Bhaskar Mitra, Emine Yilmaz +5

This is the fifth year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human-annotated training labels a…

cs.IR2025

FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents

Nandan Thakur, Jimmy Lin, Sam Havens +3

We introduce FreshStack, a holistic framework for automatically building information retrieval (IR) evaluation benchmarks by incorporating challenging questions and answers. FreshS…

cs.IR2025

Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses

Sahel Sharifymoghaddam, Shivani Upadhyay, Nandan Thakur +2

Battles, or side-by-side comparisons in so-called arenas that elicit human preferences, have emerged as a popular approach for assessing the output quality of LLMs. Recently, this…