29 papers
Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary
S. Ashwin Hebbar, Peiyao Sheng, Sewoong Oh +1
Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that…
APMM: Automated Parlay Market Maker
Niusha Moshrefi, Ranvir Rana, Pramod Viswanath
Parlays - joint contracts on the simultaneous resolution of several events - are among the most heavily traded products in betting markets, but prediction markets have struggled to…
Context-Aware RL for Agentic and Multimodal LLMs
Peiyang Xu, Bangzheng Li, Sijia Liu +4
Large language models (LLMs) often fail when answering requires identifying a small but decisive piece of evidence within a long or complex context, such as a single line in a tool…
Crypto x AI, AI x Crypto: A Survey
Sarah Allen, Pranay Anchuri, James Austgen +23
The intersection of crypto x AI is spawning papers, products, online posts, and companies. All the surrounding buzz, though, obscures what exactly has been done, what the opportuni…
TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks
Jianzhu Yao, Hongxu Su, Taobo Liao +4
Neural networks increasingly run on hardware outside the user's control (cloud GPUs, inference marketplaces). Yet ML-as-a-Service reveals little about what actually ran or whether…
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
Kevin Wang, Anna Thöni, Benjamin Kempinski +50
Large language models (LLMs) are increasingly deployed as interactive agents, yet their capacity for social and strategic reasoning over extended interaction remains poorly underst…