agent evaluation 1asynchronous runtime 1benchmarking 1large language models 1software infrastructure 1
From the 1 of 11 linked papers with an AI index.
Showing 2024Show all
2 papers · 1 filter
cs.CL2024
GTA: A Benchmark for General Tool Agents
Jize Wang, Zerun Ma, Yining Li +4
Significant focus has been placed on integrating large language models (LLMs) with various tools in developing general-purpose agents. This poses a challenge to LLMs' tool-use capa…
cs.CL2024
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
Chuyu Zhang, Songyang Zhang, Yingfan Hu +8
While LLM-Based agents, which use external tools to solve complex problems, have made significant progress, benchmarking their ability is challenging, thereby hindering a clear und…