Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games
Mingyuan Fan, Weiguang Han, Daixin Wang +3
We introduce a multi-turn interactive framework for reasoning evaluation that treats reasoning as active evidence acquisition and belief updating. Wherein, LLMs receive only the ta…
cs.AI2025
Reinforcement Learning Foundations for Deep Research Systems: A Survey
Wenjun Li, Zhi Chen, Jingru Lin +8
Deep research systems, agentic AI that solve complex, multi-step tasks by coordinating reasoning, search across the open web and user files, and tool use, are moving toward hierarc…
cs.AI2025
From Long to Short: LLMs Excel at Trimming Own Reasoning Chains
Wei Han, Geng Zhan, Sicheng Yu +2
O1/R1 style large reasoning models (LRMs) signal a substantial leap forward over conventional instruction-following LLMs. By applying test-time scaling to generate extended reasoni…