most citedA Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look

9 citations · 23 across the 6 of their papers we have counts for

collaborators

7 papers

cs.IR2025

On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools

Shivani Upadhyay, Messiah Ataey, Syed Shariyar Murtaza +2

The proliferation of complex structured data in hybrid sources, such as PDF documents and web pages, presents unique challenges for current Large Language Models (LLMs) and Multi-m…

cs.IR2025

Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses

Sahel Sharifymoghaddam, Shivani Upadhyay, Nandan Thakur +2

Battles, or side-by-side comparisons in so-called arenas that elicit human preferences, have emerged as a popular approach for assessing the output quality of LLMs. Recently, this…

cs.CL20252 cited

Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges

Nandan Thakur, Ronak Pradeep, Shivani Upadhyay +3

Retrieval-augmented generation (RAG) enables large language models (LLMs) to generate answers with citations from source documents containing "ground truth", thereby reducing syste…

cs.IR20252 cited

The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models

Ronak Pradeep, Nandan Thakur, Shivani Upadhyay +3

Large Language Models (LLMs) have significantly enhanced the capabilities of information access systems, especially with retrieval-augmented generation (RAG). Nevertheless, the eva…

cs.IR20246 cited

Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework

Ronak Pradeep, Nandan Thakur, Shivani Upadhyay +3

This report provides an initial look at partial results from the TREC 2024 Retrieval-Augmented Generation (RAG) Track. We have identified RAG evaluation as a barrier to continued p…

cs.IR20249 cited

A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look

Shivani Upadhyay, Ronak Pradeep, Nandan Thakur +5

The application of large language models to provide relevance assessments presents exciting opportunities to advance information retrieval, natural language processing, and beyond,…