Evaluating Language Model Agency through Negotiations
arXiv:2401.04536
Abstract
We introduce an approach to evaluate language model (LM) agency using negotiation games. This approach better reflects real-world use cases and addresses some of the shortcomings of alternative LM benchmarks. Negotiation games enable us to study multi-turn, and cross-model interactions, modulate complexity, and side-step accidental evaluation data leakage. We use our approach to test six widely used and publicly accessible LMs, evaluating performance and alignment in both self-play and cross-play settings. Noteworthy findings include: (i) only closed-source models tested here were able to complete these tasks; (ii) cooperative bargaining games proved to be most challenging to the models; and (iii) even the most powerful models sometimes "lose" to weaker opponents
Accepted to ICLR 2024, code and link to project data are made available at https://github.com/epfl-dlab/LAMEN
References in corpus (31)
- Training language models to follow instructions with human feedback
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- ReAct: Synergizing Reasoning and Acting in Language Models
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- QLoRA: Efficient Finetuning of Quantized LLMs
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- The Rise and Potential of Large Language Model Based Agents: A Survey
- How is ChatGPT's behavior changing over time?
- Self-Instruct: Aligning Language Models with Self-Generated Instructions
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
- Analyzing Information Leakage of Updates to Natural Language Models
- Artificial Artificial Artificial Intelligence: Crowd Workers Widely Use Large Language Models for Text Production Tasks
- ChatDev: Communicative Agents for Software Development
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
- AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
- Rethinking with Retrieval: Faithful Large Language Model Inference
- Fundamental Limitations of Alignment in Large Language Models
- Improving Language Model Negotiation with Self-Play and In-Context Learning from AI Feedback
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Reward Design with Language Models
- Reinforced Self-Training (ReST) for Language Modeling
- Prevalence and prevention of large language model use in crowd work
- GPT in Game Theory Experiments
- Thought Cloning: Learning to Think while Acting by Imitating Human Thinking
- Elo Uncovered: Robustness and Best Practices in Language Model Evaluation
- Flows: Building Blocks of Reasoning and Collaborating AI
- Generative Models as a Complex Systems Science: How can we make sense of large language model behavior?