3 papers
cs.AI2026
Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting
Andrea Wynn, Metod Jazbec, Charith Peris +4
Large language models (LLMs) can be influenced by harmful or irrelevant context, which can significantly harm model performance on downstream tasks. This motivates principled desig…
cs.CL2025
Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate
Andrea Wynn, Harsh Satija, Gillian Hadfield
While multi-agent debate has been proposed as a promising strategy for improving AI reasoning ability, we find that debate can sometimes be harmful rather than helpful. Prior work…
cs.AI2024
Learning Human-like Representations to Enable Learning Human Values
Andrea Wynn, Ilia Sucholutsky, Thomas L. Griffiths
How can we build AI systems that can learn any set of individual human values both quickly and safely, avoiding causing harm or violating societal standards for acceptable behavior…