10 citations · 24 across the 26 of their papers we have counts for
4 papers · 1 filter
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
Jielin Qiu, Zuxin Liu, Zhiwei Liu +18
As large language models (LLMs) evolve into sophisticated autonomous agents capable of complex software development tasks, evaluating their real-world capabilities becomes critical…
LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering
Jielin Qiu, Zuxin Liu, Zhiwei Liu +14
The emergence of long-context language models with context windows extending to millions of tokens has created new opportunities for sophisticated code understanding and software d…
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
Shirley Kokane, Ming Zhu, Tulika Awalgaonkar +15
Evaluating Large Language Models (LLMs) is one of the most critical aspects of building a performant compound AI system. Since the output from LLMs propagate to downstream steps, i…
Diversity Empowers Intelligence: Integrating Expertise of Software Engineering Agents
Kexun Zhang, Weiran Yao, Zuxin Liu +13
Large language model (LLM) agents have shown great potential in solving real-world software engineering (SWE) problems. The most advanced open-source SWE agent can resolve over 27%…