7 papers
Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary
S. Ashwin Hebbar, Peiyao Sheng, Sewoong Oh +1
Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that…
Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models
Kyuyoung Kim, Kevin Wang, Yunfei Xie +7
Training language models to produce both correct answers and sound reasoning remains an open challenge. Reinforcement learning with verifiable rewards typically optimizes only fina…
MURMUR: Using cross-user chatter to break collaborative language agents in groups
Atharv Singh Patlan, Peiyao Sheng, S. Ashwin Hebbar +2
Language agents are rapidly expanding from single-user assistants to multi-user collaborators in shared workspaces and groups. However, today's language models lack a mechanism for…
Scalable Fingerprinting of Large Language Models
Anshul Nasery, Jonathan Hayase, Creston Brooks +4
Model fingerprinting has emerged as a powerful tool for model owners to identify their shared model given API access. However, to lower false discovery rate, fight fingerprint leak…
Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents
Atharv Singh Patlan, Peiyao Sheng, S. Ashwin Hebbar +2
AI agents integrated with Web3 offer autonomy and openness but raise security concerns as they interact with financial protocols and immutable smart contracts. This paper investiga…
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
Zihan Zheng, Zerui Cheng, Zeyu Shen +16
Recent reports claim that large language models (LLMs) now outperform elite humans in competitive programming. Drawing on knowledge from a group of medalists in international algor…