Evaluation of Retrieval-Augmented Generation: A Survey
arXiv:2405.07437 · doi:10.1007/978-981-96-1024-2_8
Abstract
Retrieval-Augmented Generation (RAG) has recently gained traction in natural language processing. Numerous studies and real-world applications are leveraging its ability to enhance generative models through external information retrieval. Evaluating these RAG systems, however, poses unique challenges due to their hybrid structure and reliance on dynamic knowledge sources. To better understand these challenges, we conduct A Unified Evaluation Process of RAG (Auepora) and aim to provide a comprehensive overview of the evaluation and benchmarks of RAG systems. Specifically, we examine and compare several quantifiable metrics of the Retrieval and Generation components, such as relevance, accuracy, and faithfulness, within the current RAG benchmarks, encompassing the possible output and ground truth pairs. We then analyze the various datasets and metrics, discuss the limitations of current benchmarks, and suggest potential directions to advance the field of RAG benchmarks.
References in corpus (16)
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Retrieval-Augmented Generation for AI-Generated Content: A Survey
- Ragas: Automated Evaluation of Retrieval Augmented Generation
- A Survey on Retrieval-Augmented Text Generation for Large Language Models
- Benchmarking Large Language Models in Retrieval-Augmented Generation
- Benchmarking Retrieval-Augmented Generation for Medicine
- MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries
- Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
- Seven Failure Points When Engineering a Retrieval Augmented Generation System
- CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models
- ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
- RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge
- Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents
- DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
- FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation
- Let LLMs Take on the Latest Challenges! A Chinese Dynamic Question Answering Benchmark
Cited by in corpus (6)
- Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets
- SciDaSynth: Interactive Structured Data Extraction from Scientific Literature with Large Language Model
- FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis
- SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots
- Retrieval-Augmented Generation for Natural Language Art Provenance Searches in the Getty Provenance Index
- A Retrieval-Augmented Automated Stakeholder for Requirements Elicitation Education: A Comparative Study