1 paper
Yuyang Wu, Yue Huang, Shuaike Shen +8
Large Language Models (LLMs) have become increasingly capable as tool-using agents, with benchmarks spanning diverse general agentic tasks. Yet rigorous evaluation of scientific to…