37 citations · 38 across the 17 of their papers we have counts for
Showing cs.SEShow all
3 papers · 1 filter
cs.SE2026
Understanding by Reconstruction: Reversing the Software Development Process for LLM Pretraining
Zhiyuan Zeng, Yichi Zhang, Yong Shan +11
While Large Language Models (LLMs) have achieved remarkable success in code generation, they often struggle with the deep, long-horizon reasoning required for complex software engi…
cs.SE2026
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
Yubang Wang, Chenxi Zhang, Bowen Chen +7
Autonomous agents are increasingly expected to support scientific research, and recent benchmarks report progress in code repair and autonomous experimentation. However, these eval…
cs.SE2026
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
Jie Yang, Honglin Guo, Li Ji +11
The evolution of Large Language Models (LLMs) into autonomous agents has expanded the scope of AI coding from localized code generation to complex, repository-level, and execution-…