Showing cs.SEShow all
2 papers · 1 filter
cs.SE2025
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
Jielin Qiu, Zuxin Liu, Zhiwei Liu +18
As large language models (LLMs) evolve into sophisticated autonomous agents capable of complex software development tasks, evaluating their real-world capabilities becomes critical…
cs.SE2025
LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering
Jielin Qiu, Zuxin Liu, Zhiwei Liu +14
The emergence of long-context language models with context windows extending to millions of tokens has created new opportunities for sophisticated code understanding and software d…