Showing cs.SEShow all
2 papers · 1 filter
cs.SE2026
DeepDiscovery: A Location-Inference Framework for Task-Level Repository Understanding
Jiawei He, Weisong Sun, Mengyu Shi +4
Large language models have shown strong performance on software engineering (SE) tasks, yet understanding large industrial repositories remains challenging. Existing methods often…
cs.SE2026
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
Jiawei He, Jie Jia, Chenbo Liu +4
Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limited visibility and often miss…