Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
cs.AI2026
Enhancing Table Reasoning with Deterministic Table-State Rewards
Tung Sum Thomas Kwok, Xinyu Wang, Hengzhi He +9
Large Language Models (LLMs) struggle with multi-step reasoning over structured tables. The primary reason is the lack of explicit supervision for intermediate reasoning states. Ex…