3 papers
cs.SE2026
On the Effectiveness of Context Compression for Repository-Level Tasks: An Empirical Investigation
Jia Feng, Zhanyue Qin, Cuiyun Gao +4
Repository-level code intelligence tasks require large language models (LLMs) to process long, multi-file contexts. Such inputs introduce three challenges: crucial context can be o…
cs.AI2026
DockSmith: Scaling Reliable Coding Environments via an Agentic Docker Builder
Jiaran Zhang, Luck Ma, Fanqi Wan +10
Reliable Docker-based environment construction is a dominant bottleneck for scaling execution-grounded training and evaluation of software engineering agents. We introduce DockSmit…
cs.SE2025
Multi-Docker-Eval: A `Shovel of the Gold Rush' Benchmark on Automatic Environment Building for Software Engineering
Kelin Fu, Tianyu Liu, Zeyu Shang +4
Automated environment configuration is a critical bottleneck in scaling software engineering (SWE) automation. To provide a reliable evaluation standard for this task, we present M…