Tool Learning with Large Language Models: A Survey
arXiv:2405.17935 · doi:10.1007/s11704-024-40678-2
Abstract
Recently, tool learning with large language models (LLMs) has emerged as a promising paradigm for augmenting the capabilities of LLMs to tackle highly complex problems. Despite growing attention and rapid advancements in this field, the existing literature remains fragmented and lacks systematic organization, posing barriers to entry for newcomers. This gap motivates us to conduct a comprehensive survey of existing works on tool learning with LLMs. In this survey, we focus on reviewing existing literature from the two primary aspects (1) why tool learning is beneficial and (2) how tool learning is implemented, enabling a comprehensive understanding of tool learning with LLMs. We first explore the "why" by reviewing both the benefits of tool integration and the inherent benefits of the tool learning paradigm from six specific aspects. In terms of "how", we systematically review the literature according to a taxonomy of four key stages in the tool learning workflow: task planning, tool selection, tool calling, and response generation. Additionally, we provide a detailed summary of existing benchmarks and evaluation methods, categorizing them according to their relevance to different stages. Finally, we discuss current challenges and outline potential future directions, aiming to inspire both researchers and industrial developers to further explore this emerging and promising area. We also maintain a GitHub repository to continually keep track of the relevant papers and resources in this rising area at https://github.com/quchangle1/LLM-Tool-Survey.
The article has been accepted by Frontiers of Computer Science (FCS), with the DOI: {10.1007/s11704-024-40678-2}
References in corpus (67)
- Survey of Hallucination in Natural Language Generation
- A Survey on Large Language Model based Autonomous Agents
- A Survey of Large Language Models
- Evaluating Large Language Models Trained on Code
- A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly
- Gemini: A Family of Highly Capable Multimodal Models
- Retrieval-Augmented Generation for Large Language Models: A Survey
- LaMDA: Language Models for Dialog Applications
- The Rise and Potential of Large Language Model Based Agents: A Survey
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
- GeneGPT: Augmenting Large Language Models with Domain Tools for Improved Access to Biomedical Information
- Gorilla: Large Language Model Connected with Massive APIs
- MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
- C-Pack: Packed Resources For General Chinese Embeddings
- Ethical and social risks of harm from Language Models
- Internet-augmented language models through few-shot prompting for open-domain question answering
- Investigating the Effectiveness of ChatGPT in Mathematical Reasoning and Problem Solving: Evidence from the Vietnamese National High School Graduation Examination
- ART: Automatic multi-step reasoning and tool-use for large language models
- MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning
- Let Me Do It For You: Towards LLM Empowered Recommendation via Tool Learning
- Understanding the planning of LLM agents: A survey
- Tool Learning with Foundation Models
- TALM: Tool Augmented Language Models
- Security and Privacy Challenges of Large Language Models: A Survey
- FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios
- AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn
- A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
- ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
- Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems
- Graph-ToolFormer: To Empower LLMs with Graph Reasoning Ability via Prompt Augmented by ChatGPT
- On the Tool Manipulation Capability of Open-source Large Language Models
- ToolCoder: Teach Code Generation Models to use API search tools
- Tool Documentation Enables Zero-Shot Tool-Usage with Large Language Models
- CMATH: Can Your Language Model Pass Chinese Elementary School Math Test?
- Towards Completeness-Oriented Tool Retrieval for Large Language Models
- TaskBench: Benchmarking Large Language Models for Task Automation
- TPTU-v2: Boosting Task Planning and Tool Usage of Large Language Model-based Agents in Real-world Systems
- Tool Learning in the Wild: Empowering Language Models as Automatic Tool Agents
- ControlLLM: Augment Language Models with Tools by Searching on Graphs
- Efficient Tool Use with Chain-of-Abstraction Reasoning
- STRIDE: A Tool-Assisted LLM Agent Framework for Strategic and Interactive Decision-Making
- ToolNet: Connecting Large Language Models with Massive Tools via Tool Graph
- EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction
- TPE: Towards Better Compositional Reasoning over Conceptual Tools with Multi-persona Collaboration
- Towards Verifiable Text Generation with Evolving Memory and Self-Reflection
- Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub
- From Summary to Action: Enhancing Large Language Models for Complex Tasks with Open World APIs
- Simulating Financial Market via Large Language Model based Agents
- ToolTalk: Evaluating Tool-Usage in a Conversational Setting
- GTA: A Benchmark for General Tool Agents
- Can Graph Learning Improve Planning in LLM-based Agents?
- ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
- m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
- TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement
- APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
- Don't Fine-Tune, Decode: Syntax Error-Free Tool Use via Constrained Decoding
- MathViz-E: A Case-study in Domain-Specialized Tool-Using Agents
- From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions
- Look Before You Leap: Towards Decision-Aware and Generalizable Tool-Usage for Large Language Models
- WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models
- Evaluating Tool-Augmented Agents in Remote Sensing Platforms
- ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios
- ProTIP: Progressive Tool Retrieval Improves Planning
- VIoTGPT: Learning to Schedule Vision Tools in LLMs towards Intelligent Video Internet of Things
- Tool-Planner: Task Planning with Clusters across Multiple Tools
- Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference Trees
- ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents