2 papers
cs.AI2026
What Process Evaluation of Coding Agents Actually Measures: Action, Task, and Step Are Three Different Levels
Jiawei He, Mengyu Shi, Jie jia +2
Coding agents are increasingly evaluated not only by whether they solve a task, but also by how they execute it. However, existing process-level evaluations often treat action pred…
cs.SE2026
DeepDiscovery: A Location-Inference Framework for Task-Level Repository Understanding
Jiawei He, Weisong Sun, Mengyu Shi +4
Large language models have shown strong performance on software engineering (SE) tasks, yet understanding large industrial repositories remains challenging. Existing methods often…