RELIC: Investigating Large Language Model Responses using Self-Consistency
arXiv:2311.16842 · doi:10.1145/3613904.3641904
Abstract
Large Language Models (LLMs) are notorious for blending fact with fiction and generating non-factual content, known as hallucinations. To address this challenge, we propose an interactive system that helps users gain insight into the reliability of the generated text. Our approach is based on the idea that the self-consistency of multiple samples generated by the same LLM relates to its confidence in individual claims in the generated texts. Using this idea, we design RELIC, an interactive system that enables users to investigate and verify semantic-level variations in multiple long-form responses. This allows users to recognize potentially inaccurate information in the generated text and make necessary corrections. From a user study with ten participants, we demonstrate that our approach helps users better verify the reliability of the generated text. We further summarize the design implications and lessons learned from this research for future studies of reliable human-LLM interactions.
References in corpus (10)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model Capabilities
- explAIner: A Visual Analytics Framework for Interactive and Explainable Machine Learning
- Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language Models
- Graphologue: Exploring Large Language Model Responses with Interactive Diagrams
- Beyond Text Generation: Supporting Writers with Continuous Automatic Text Summaries
- The State of Human-centered NLP Technology for Fact-checking
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Visual Comparison of Language Model Adaptation
Cited by in corpus (5)
- RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented Generation
- Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions
- DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition
- Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research
- Visual Text Mining with Progressive Taxonomy Construction for Environmental Studies