collaborators

8 papers

cs.SE2026

LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation

Han Li, Zhemin Fang, Rili Feng +8

Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon software development. Existing ben…

cs.SE2026

Failure as a Process: An Anatomy of CLI Coding Agent Trajectories

Xiangxin Zhao, Han Li, Shuaiting Li +4

Large language model (LLM) coding agents are increasingly deployed to autonomously perform software engineering tasks in terminal-based environments, making their reliability a gro…

cs.SE2026

An Iterative Test-and-Repair Framework for Competitive Code Generation

Lingxiao Tang, Muyang Ye, Zhaoyang Chu +4

Large language models (LLMs) have made remarkable progress in code generation, but competitive programming remains a challenge. Recent training-based methods have improved code gen…

cs.AI2026

TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks

Zhaoyang Chu, Jiarui Hu, Xingyu Jiang +8

We introduce TerminalWorld, a scalable data engine that automatically reverse-engineers high-fidelity evaluation tasks from "in-the-wild" terminal recordings. Processing 80,870 ter…

cs.SE2026

HerAgent: Rethinking the Automated Environment Deployment via Hierarchical Test Pyramid

Xiang Li, Siyu Lu, Federica Sarro +2

Automated software environment setup is a prerequisite for testing, debugging, and reproducing failures, yet remains challenging in practice due to complex dependencies, heterogene…

cs.LG2026

ContextBench: A Benchmark for Context Retrieval in Coding Agents

Han Li, Letian Zhu, Bohan Zhang +7

LLM-based coding agents have shown strong performance on automated issue resolution benchmarks, yet existing evaluations largely focus on final task success, providing limited insi…