Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary
S. Ashwin Hebbar, Peiyao Sheng, Sewoong Oh +1
Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that…
cs.CL2026
Context-Aware RL for Agentic and Multimodal LLMs
Peiyang Xu, Bangzheng Li, Sijia Liu +4
Large language models (LLMs) often fail when answering requires identifying a small but decisive piece of evidence within a long or complex context, such as a single line in a tool…
cs.CL2026
Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models
Kyuyoung Kim, Kevin Wang, Yunfei Xie +7
Training language models to produce both correct answers and sound reasoning remains an open challenge. Reinforcement learning with verifiable rewards typically optimizes only fina…