37 citations · 68 across the 11 of their papers we have counts for
16 papers · 1 filter
GRAFITE: Generative Regression Analysis Framework for Issue Tracking and Evaluation
Ja Young Lee, Mírian Silva, Mohamed Nasr +6
Large language models (LLMs) are largely motivated by their performance on popular topics and benchmarks at the time of their release. However, over time, contamination occurs due…
MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations
Sara Rosenthal, Yannis Katsis, Vraj Shah +3
We present MTRAG-UN, a benchmark for exploring open challenges in multi-turn retrieval augmented generation, a popular use of large language models. We release a benchmark of 666 t…
RAGAPHENE: A RAG Annotation Platform with Human Enhancements and Edits
Kshitij Fadnis, Sara Rosenthal, Maeda Hanafi +2
Retrieval Augmented Generation (RAG) is an important aspect of conversing with Large Language Models (LLMs) when factually correct information is important. LLMs may provide answer…
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
Yannis Katsis, Sara Rosenthal, Kshitij Fadnis +7
Retrieval-augmented generation (RAG) has recently become a very popular task for Large Language Models (LLMs). Evaluating them on multi-turn RAG conversations, where the system is…
CLAPNQ: Cohesive Long-form Answers from Passages in Natural Questions for RAG systems
Sara Rosenthal, Avirup Sil, Radu Florian +1
Retrieval Augmented Generation (RAG) has become a popular application for large language models. It is preferable that successful RAG systems provide accurate answers that are supp…
Muted: Multilingual Targeted Offensive Speech Identification and Visualization
Christoph Tillmann, Aashka Trivedi, Sara Rosenthal +4
Offensive language such as hate, abuse, and profanity (HAP) occurs in various content on the web. While previous work has mostly dealt with sentence level annotations, there have b…