Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
Jacob Eisenstein, Fantine Huot, Adam Fisch +2
We present a scalable and verifiable methodology for evaluating language models in multi-turn interactions, using a suite of collaborative games that require effective communicatio…
cs.CL2025
Comparing Human and Language Models Sentence Processing Difficulties on Complex Structures
Samuel Joseph Amouyal, Aya Meltzer-Asscher, Jonathan Berant
Large language models (LLMs) that fluently converse with humans are a reality - but do LLMs experience human-like processing difficulties? We systematically compare human and LLM s…