3 papers
cs.SE2026
NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs
Jiarong Zhao, Zhikai Lei, Zhiheng Xi +5
The paper presents NexForge, a requirement‑first framework that automatically turns free‑form capability requirements into executable agent training tasks, scaling data generation…
cs.AI2026
Steering LLMs via Scalable Interactive Oversight
Enyu Zhou, Zhiheng Xi, Long Ma +9
As Large Language Models increasingly automate complex, long-horizon tasks such as \emph{vibe coding}, a supervision gap has emerged. While models excel at execution, users often s…
cs.SE2026
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
Jie Yang, Honglin Guo, Li Ji +11
The evolution of Large Language Models (LLMs) into autonomous agents has expanded the scope of AI coding from localized code generation to complex, repository-level, and execution-…