ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing
arXiv:2309.09128 · doi:10.1145/3613904.3642016
Abstract
Evaluating outputs of large language models (LLMs) is challenging, requiring making -- and making sense of -- many responses. Yet tools that go beyond basic prompting tend to require knowledge of programming APIs, focus on narrow domains, or are closed-source. We present ChainForge, an open-source visual toolkit for prompt engineering and on-demand hypothesis testing of text generation LLMs. ChainForge provides a graphical interface for comparison of responses across models and prompt variations. Our system was designed to support three tasks: model selection, prompt template design, and hypothesis testing (e.g., auditing). We released ChainForge early in its development and iterated on its design with academics and online users. Through in-lab and interview studies, we find that a range of people could use ChainForge to investigate hypotheses that matter to them, including in real-world settings. We identify three modes of prompt engineering and LLM hypothesis testing: opportunistic exploration, limited evaluation, and iterative refinement.
18 pages, 7 figures, published at CHI 2024
References in corpus (4)
- Designing for Responsible Trust in AI Systems: A Communication Perspective
- Graphologue: Exploring Large Language Model Responses with Interactive Diagrams
- Prompting Is Programming: A Query Language for Large Language Models
- Model Sketching: Centering Concepts in Early-Stage Machine Learning Model Design
Cited by in corpus (22)
- DirectGPT: A Direct Manipulation Interface to Interact with Large Language Models
- Generative AI for Self-Adaptive Systems: State of the Art and Research Roadmap
- Assessing the Ability of ChatGPT to Screen Articles for Systematic Reviews
- What Should We Engineer in Prompts? Training Humans in Requirement-Driven LLM Use
- Interactive Debugging and Steering of Multi-Agent AI Systems
- Patchview: LLM-Powered Worldbuilding with Generative Dust and Magnet Visualization
- Towards Dataset-scale and Feature-oriented Evaluation of Text Summarization in Large Language Model Prompts
- ChainBuddy: An AI Agent System for Generating LLM Pipelines
- SPROUT: an Interactive Authoring Tool for Generating Programming Tutorials with the Visualization of Large Language Models
- Access Denied: Meaningful Data Access for Quantitative Algorithm Audits
- Interrogating AI: Characterizing Emergent Playful Interactions with ChatGPT
- Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are Absent
- InstructPipe: Generating Visual Blocks Pipelines with Human Instructions and LLMs
- RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented Generation
- Understanding the Dataset Practitioners Behind Large Language Model Development
- OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models
- Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions
- On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards
- Vipera: Towards systematic auditing of generative text-to-image models at scale
- Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
- DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition
- OOPrompt: Reifying Intents into Structured Artifacts for Modular and Iterative Prompting