Evaluation of Retrieval-Augmented Generation: A Survey
arXiv:2405.07437 · doi:10.1007/978-981-96-1024-2_8
Abstract
Retrieval-Augmented Generation (RAG) has recently gained traction in natural language processing. Numerous studies and real-world applications are leveraging its ability to enhance generative models through external information retrieval. Evaluating these RAG systems, however, poses unique challenges due to their hybrid structure and reliance on dynamic knowledge sources. To better understand these challenges, we conduct A Unified Evaluation Process of RAG (Auepora) and aim to provide a comprehensive overview of the evaluation and benchmarks of RAG systems. Specifically, we examine and compare several quantifiable metrics of the Retrieval and Generation components, such as relevance, accuracy, and faithfulness, within the current RAG benchmarks, encompassing the possible output and ground truth pairs. We then analyze the various datasets and metrics, discuss the limitations of current benchmarks, and suggest potential directions to advance the field of RAG benchmarks.
References in corpus (25)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Emergent Abilities of Large Language Models
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- ReCoRD: Bridging the Gap between Human and Machine Commonsense Reading Comprehension
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
- Retrieval-Augmented Generation for AI-Generated Content: A Survey
- Ragas: Automated Evaluation of Retrieval Augmented Generation
- Incorporating Dynamic Semantics into Pre-Trained Language Model for Aspect-based Sentiment Analysis
- A Survey on Retrieval-Augmented Text Generation for Large Language Models
- Benchmarking Large Language Models in Retrieval-Augmented Generation
- RealTime QA: What's the Answer Right Now?
- Benchmarking Retrieval-Augmented Generation for Medicine
- MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries
- Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
- Seven Failure Points When Engineering a Retrieval Augmented Generation System
- CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models
- Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents
- ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
- RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge
- DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
- Learning to Reason and Memorize with Self-Notes
- FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation
- Let LLMs Take on the Latest Challenges! A Chinese Dynamic Question Answering Benchmark
Cited by in corpus (6)
- Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets
- SciDaSynth: Interactive Structured Data Extraction from Scientific Literature with Large Language Model
- FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis
- SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots
- Retrieval-Augmented Generation for Natural Language Art Provenance Searches in the Getty Provenance Index
- A Retrieval-Augmented Automated Stakeholder for Requirements Elicitation Education: A Comparative Study