ChainBuddy: An AI Agent System for Generating LLM Pipelines
arXiv:2409.13588 · doi:10.1145/3706598.3714085
Abstract
As large language models (LLMs) advance, their potential applications have grown significantly. However, it remains difficult to evaluate LLM behavior on user-defined tasks and craft effective pipelines to do so. Many users struggle with where to start, often referred to as the "blank page problem." ChainBuddy, an AI workflow generation assistant built into the ChainForge platform, aims to tackle this issue. From a single prompt or chat, ChainBuddy generates a starter evaluative LLM pipeline in ChainForge aligned to the user's requirements. ChainBuddy offers a straightforward and user-friendly way to plan and evaluate LLM behavior and make the process less daunting and more accessible across a wide range of possible tasks and use cases. We report a within-subjects user study comparing ChainBuddy to the baseline interface. We find that when using AI assistance, participants reported a less demanding workload, felt more confident, and produced higher quality pipelines evaluating LLM behavior. However, we also uncover a mismatch between subjective and objective ratings of performance: participants rated their successfulness similarly across conditions, while independent experts rated participant workflows significantly higher with AI assistance. Drawing connections to the Dunning-Kruger effect, we draw design implications for the future of workflow generation assistants to mitigate the risk of over-reliance.
21 pages, 12 figures, pre-print
References in corpus (10)
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- Telling Stories from Computational Notebooks: AI-Assisted Presentation Slides Creation for Presenting Data Science Work
- Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition
- DSPy Assertions: Computational Constraints for Self-Refining Language Model Pipelines
- AutoEval Done Right: Using Synthetic Data for Model Evaluation
- Aligning Language Models with Demonstrated Feedback