1 paper
Genta Indra Winata, Amartya Chakraborty, Yuzhen Lin +12
Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However, existing benchmarks remain li…