Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems
Shunwen Bai, Ziping Ma, Chaoyang Zhang +4
The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and all…
cs.AI2024
Infant Agent: A Tool-Integrated, Logic-Driven Agent with Cost-Effective API Usage
Bin Lei, Yuchen Li, Yiming Zeng +7
Despite the impressive capabilities of large language models (LLMs), they currently exhibit two primary limitations, \textbf{\uppercase\expandafter{\romannumeral 1}}: They struggle…