Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability
Yang Tian, Zhengpeng Shi, Yu Zhou +1
Large language models are increasingly deployed as agents that solve tasks by interacting with external tool environments. Although recent tool-use benchmarks increasingly cover co…
cs.CL2025
MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers
Yang Tian, Zheng Lu, Mingqi Gao +2
Fully comprehending scientific papers by machines reflects a high level of Artificial General Intelligence, requiring the ability to reason across fragmented and heterogeneous sour…