3 papers
cs.SE2026
Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?
Silin Chen, Yufei Yang, Xiaodong Gu +3
Repository-level coding benchmarks have become the standard for evaluating coding agents, yet they inherently suffer from data leakage because they are built upon popular open-sour…
cs.SE2026
Repo0: Design-Driven Zero-to-All Code Generation
Silin Chen, Haoyi Teng, Xiaodong Gu +5
Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold…
cs.SE2026
SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution
Silin Chen, Han Li, Xiaodong Gu +2
Large language model (LLM) based agents have demonstrated remarkable proficiency in automated software issue resolution, yet they often struggle to resolve issues in a specific rep…