5 papers
Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems
Mason Nakamura, Abhinav Kumar, Saswat Das +5
Multi-agent systems, where LLM agents communicate through free-form language, enable sophisticated coordination for solving complex cooperative tasks. This surfaces a unique safety…
Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models
Saaduddin Mahmud, Mason Nakamura, Kyle Hollins Wray +1
Prompt optimization methods have demonstrated significant effectiveness in aligning black-box large language models (LLMs). In parallel, inference scaling strategies such as Best-o…
Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
Mason Nakamura, Abhinav Kumar, Saaduddin Mahmud +3
A multi-agent system (MAS) powered by large language models (LLMs) can automate tedious user tasks such as meeting scheduling that requires inter-agent collaboration. LLMs enable n…
Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
Mason Nakamura, Saaduddin Mahmud, Kyle H. Wray +2
Aligning LLMs with user preferences is crucial for real-world use but often requires costly fine-tuning or expensive inference, forcing trade-offs between alignment quality and com…
MAPLE: A Framework for Active Preference Learning Guided by Large Language Models
Saaduddin Mahmud, Mason Nakamura, Shlomo Zilberstein
The advent of large language models (LLMs) has sparked significant interest in using natural language for preference learning. However, existing methods often suffer from high comp…