Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Agents Trust Tools Too Much: Measuring Reliance on Unreliable Tools
Hoyeol Yang, Woojung Song, Taewon Kim +3
Existing evaluations of tool-using agents primarily measure whether an agent can successfully complete diverse tasks with tools. These evaluations generally assume that tools retur…
cs.AI2026
SHAPE of Chain-of-Thought in Math Reasoning
Jonghyun Song, Sangjun Song, Minjae Oh +3
Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically meaningful skills underlying their reasoning remain underexplor…