OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models
arXiv:2508.21061 · doi:10.1145/3746059.3747746
Abstract
As multi-turn dialogues with large language models (LLMs) grow longer and more complex, how can users better evaluate and review progress on their conversational goals? We present OnGoal, an LLM chat interface that helps users better manage goal progress. OnGoal provides real-time feedback on goal alignment through LLM-assisted evaluation, explanations for evaluation results with examples, and overviews of goal progression over time, enabling users to navigate complex dialogues more effectively. Through a study with 20 participants on a writing task, we evaluate OnGoal against a baseline chat interface without goal tracking. Using OnGoal, participants spent less time and effort to achieve their goals while exploring new prompting strategies to overcome miscommunication, suggesting tracking and visualizing goals can enhance engagement and resilience in LLM dialogues. Our findings inspired design implications for future LLM chat interfaces that improve goal communication, reduce cognitive load, enhance interactivity, and enable feedback to improve LLM performance.
Accepted to UIST 2025. 18 pages, 9 figures, 2 tables. For a demo video, see https://youtu.be/uobhmxo6EIE
References in corpus (12)
- The Programmer's Assistant: Conversational Interaction with a Large Language Model for Software Development
- The Metacognitive Demands and Opportunities of Generative AI
- Survey on Evaluation Methods for Dialogue Systems
- Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language Models
- ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing
- Graphologue: Exploring Large Language Model Responses with Interactive Diagrams
- DirectGPT: A Direct Manipulation Interface to Interact with Large Language Models
- A Taxonomy for Human-LLM Interaction Modes: An Initial Exploration
- Understanding Users' Dissatisfaction with ChatGPT Responses: Types, Resolving Tactics, and the Effect of Knowledge Level
- Task Supportive and Personalized Human-Large Language Model Interaction: A User Study
- ScatterShot: Interactive In-context Example Curation for Text Transformation
- Developing a Conversational Recommendation System for Navigating Limited Options