3 papers
cs.CL2026
Last Translation Benchmark
Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…
cs.SE2026
Where Accountability Lives: Mapping Human Responsibility to Workflow Artifacts in Agentic Software Development
Sabry E. Farrag
Coding agents author commits, open pull requests, and push code in production repositories. Who is accountable is settled in two places that do not refer to each other: the platfor…
cs.SE2026
The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development
Sabry E. Farrag
Since 2022, AI-powered coding assistants have produced contradictory evidence: controlled studies report 20-56% productivity gains on well-scoped tasks, while the most rigorous RCT…