Analyzing the Presentation, Content, and Utilization of References in LLM-powered Conversational AI Systems
arXiv:2604.15326 · doi:10.1145/3772363.3798415
Abstract
As conversational AI systems become popular for information retrieval and question-answering, the references they cite are key to ensuring their answers are reliable and trustworthy. Yet, no prior work systematically analyzes how these references are presented or their quality. We examine 1,517 references from 30 question-answer pairs across nine systems, focusing on their (1) presentation in the user interface and (2) quality using the CRAAP criteria. We find notable variations in the presentation, quality, and quantity of references across systems. For instance, ChatGPT provides more references (9.5 per response on average) with higher quality (15.48/20 CRAAP score), while Hunyuan-TurboS provides fewer references (4.0) and lower quality (11.65/20). Additionally, a preliminary user study shows that people rarely interact with these references and that their behavior differs across systems. These findings highlight the need for better interface designs that help users engage with and trust references more effectively.
8 pages, 5 figures, Accepted to ACM CHI 2026 Extended Abstract/Poster/Case Study Track
References in corpus (19)
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making
- How is ChatGPT's behavior changing over time?
- Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking
- Impacts of Personal Characteristics on User Trust in Conversational Recommender Systems
- The simulation of judgment in LLMs
- AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap
- A Browser Extension for in-place Signaling and Assessment of Misinformation
- A Survey on Automatic Credibility Assessment Using Textual Credibility Signals in the Era of Large Language Models
- CUI@CHI 2024: Building Trust in CUIs-From Design to Deployment
- Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
- Large Language Models Reflect Human Citation Patterns with a Heightened Citation Bias
- Experience with GitHub Copilot for Developer Productivity at Zoominfo
- Guidance Source Matters: How Guidance from AI, Expert, or a Group of Analysts Impacts Visual Data Preparation and Analysis
- vitaLITy 2: Reviewing Academic Literature Using Large Language Models
- Who Gets Cited? Gender- and Majority-Bias in LLM-Driven Reference Selection
- Agentic Enterprise: AI-Centric User to User-Centric AI
- Not All Transparency Is Equal: Source Presentation Effects on Attention, Interaction, and Persuasion in Conversational Search
- When AI Persuades: Adversarial Explanation Attacks on Human Trust in AI-Assisted Decision Making
- "Always Nice and Confident, Sometimes Wrong": Developer's Experiences Engaging Large Language Models (LLMs) Versus Human-Powered Q&A Platforms for Coding Support