7 papers
CollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion
Adam Fisch, Daniel Deutsch, Joshua Maynez +5
Evaluating generative AI models is a routine, but resource-intensive, process that is conducted over and over again during the course of model development. In this work, we propose…
MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
Jacob Eisenstein, Fantine Huot, Adam Fisch +2
We present a scalable and verifiable methodology for evaluating language models in multi-turn interactions, using a suite of collaborative games that require effective communicatio…
Learning Steerable Clarification Policies with Collaborative Self-play
Jonathan Berant, Maximillian Chen, Adam Fisch +4
To handle underspecified or ambiguous queries, AI assistants need a policy for managing their uncertainty to determine (a) when to guess the user intent and answer directly, (b) wh…
Plantain: Plan-Answer Interleaved Reasoning
Anthony Liang, Jonathan Berant, Adam Fisch +3
Reasoning models often spend a significant amount of time thinking before they generate a visible response. In the meantime, they do not give the user any hints as to whether their…
Comparing Human and Language Models Sentence Processing Difficulties on Complex Structures
Samuel Joseph Amouyal, Aya Meltzer-Asscher, Jonathan Berant
Large language models (LLMs) that fluently converse with humans are a reality - but do LLMs experience human-like processing difficulties? We systematically compare human and LLM s…
Don't lie to your friends: Learning what you know from collaborative self-play
Jacob Eisenstein, Reza Aghajani, Adam Fisch +5
To be helpful assistants, AI agents must be aware of their own capabilities and limitations. This includes knowing when to answer from parametric knowledge versus using tools, when…