Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey)
arXiv:2407.12858 · doi:10.1145/3637528.3671467
Abstract
With the ongoing rapid adoption of Artificial Intelligence (AI)-based systems in high-stakes domains, ensuring the trustworthiness, safety, and observability of these systems has become crucial. It is essential to evaluate and monitor AI systems not only for accuracy and quality-related metrics but also for robustness, bias, security, interpretability, and other responsible AI dimensions. We focus on large language models (LLMs) and other generative AI models, which present additional challenges such as hallucinations, harmful and manipulative content, and copyright infringement. In this survey article accompanying our KDD 2024 tutorial, we highlight a wide range of harms associated with generative AI systems, and survey state of the art approaches (along with open challenges) to address these harms.
Survey Article for the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD 2024) Tutorial
References in corpus (12)
- Deep Learning with Differential Privacy
- Survey of Hallucination in Natural Language Generation
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- Word Embeddings Quantify 100 Years of Gender and Ethnic Stereotypes
- Bias in Bios: A Case Study of Semantic Representation Bias in a High-Stakes Setting
- Locating and Editing Factual Associations in GPT
- Rethinking Search: Making Domain Experts out of Dilettantes
- Contrastive Language-Vision AI Models Pretrained on Web-Scraped Multimodal Data Exhibit Sexual Objectification Bias
- Iterative Methods for Private Synthetic Data: Unifying Framework and New Methods
- Are Two Heads the Same as One? Identifying Disparate Treatment in Fair Neural Networks