2 papers
cs.CL2026
Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary
S. Ashwin Hebbar, Peiyao Sheng, Sewoong Oh +1
Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that…
cs.CR2025
MURMUR: Using cross-user chatter to break collaborative language agents in groups
Atharv Singh Patlan, Peiyao Sheng, S. Ashwin Hebbar +2
Language agents are rapidly expanding from single-user assistants to multi-user collaborators in shared workspaces and groups. However, today's language models lack a mechanism for…