29 citations · 43 across the 29 of their papers we have counts for
Showing cs.SEShow all
2 papers · 1 filter
cs.SE2026
SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering
Qingnan Ren, Shun Zou, Shiting Huang +11
As autonomous coding agents become capable of handling increasingly long-horizon tasks, they have gradually demonstrated the potential to complete end-to-end software development.…
cs.SE2026
CriterAlign: Criterion-Centric Rationale Alignment for Code Preference Judging
Zhenyu Li, Aleksandar Cvejic, Zehui Chen +1
Pairwise human preference prediction is central to evaluating code-generation systems, where quality often depends on task-specific trade-offs beyond functional correctness. While…