activity
20242026
collaborators

8 papers

cs.CL2026

GRAFITE: Generative Regression Analysis Framework for Issue Tracking and Evaluation

Ja Young Lee, Mírian Silva, Mohamed Nasr +6

Large language models (LLMs) are largely motivated by their performance on popular topics and benchmarks at the time of their release. However, over time, contamination occurs due…

cs.CL2026

MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations

Sara Rosenthal, Yannis Katsis, Vraj Shah +3

We present MTRAG-UN, a benchmark for exploring open challenges in multi-turn retrieval augmented generation, a popular use of large language models. We release a benchmark of 666 t…

cs.HC2025

A Longitudinal Study on Different Annotator Feedback Loops in Complex RAG Tasks

Sara Rosenthal, Maeda Hanafi, Yannis Katsis +2

Grounding conversations in existing passages, known as Retrieval-Augmented Generation (RAG), is an important aspect of Chat-Based Assistants powered by Large Language Models (LLMs)…

cs.CL2025

RAGAPHENE: A RAG Annotation Platform with Human Enhancements and Edits

Kshitij Fadnis, Sara Rosenthal, Maeda Hanafi +2

Retrieval Augmented Generation (RAG) is an important aspect of conversing with Large Language Models (LLMs) when factually correct information is important. LLMs may provide answer…

cs.SE2025

InspectorRAGet: An Introspection Platform for RAG Evaluation

Kshitij Fadnis, Siva Sankalp Patel, Odellia Boni +4

Large Language Models (LLM) have become a popular approach for implementing Retrieval Augmented Generation (RAG) systems, and a significant amount of effort has been spent on build…

cs.IR2025

Granite Embedding Models

Parul Awasthy, Aashka Trivedi, Yulong Li +19

We introduce the Granite Embedding models, a family of encoder-based embedding models designed for retrieval tasks, spanning dense-retrieval and sparse retrieval architectures, wit…