Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate
arXiv:2305.11595 · doi:10.18653/v1/2023.findings-emnlp.508
Abstract
Large Language Models (LLMs) have shown impressive capabilities in various applications, but they still face various inconsistency issues. Existing works primarily focus on the inconsistency issues within a single LLM, while we complementarily explore the inter-consistency among multiple LLMs for collaboration. To examine whether LLMs can collaborate effectively to achieve a consensus for a shared goal, we focus on commonsense reasoning, and introduce a formal debate framework (FORD) to conduct a three-stage debate among LLMs with real-world scenarios alignment: fair debate, mismatched debate, and roundtable debate. Through extensive experiments on various datasets, LLMs can effectively collaborate to reach a consensus despite noticeable inter-inconsistencies, but imbalances in their abilities can lead to domination by superior LLMs. Leveraging a more advanced LLM like GPT-4 as an authoritative judge can boost collaboration performance. Our work contributes to understanding the inter-consistency among LLMs and lays the foundation for developing future collaboration methods. Codes and data are available at https://github.com/Waste-Wood/FORD
EMNLP 2023 Findings Camera Ready Version
References in corpus (15)
- Training language models to follow instructions with human feedback
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Scaling Instruction-Finetuned Language Models
- Large Language Models are Zero-Shot Reasoners
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- LaMDA: Language Models for Dialog Applications
- Multitask Prompted Training Enables Zero-Shot Task Generalization
- Automatic Chain of Thought Prompting in Large Language Models
- Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
- Complexity-Based Prompting for Multi-Step Reasoning
- PEER: A Collaborative Language Model
- Making Large Language Models Better Reasoners with Step-Aware Verifier
- Unpacking Large Language Models with Conceptual Consistency