2 papers
cs.LG2026
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
Adarsh Kumarappan, Pareesa Ameneh Golnari, Wen Wen +5
DevBench is a telemetry-driven benchmark designed to evaluate Large Language Models (LLMs) on realistic code completion tasks. It includes 1,800 evaluation instances across six pro…
cs.SE2026
The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot
Fangchen Song, Ashish Agarwal, Wen Wen
Generative artificial intelligence (AI) facilitates content production and enhances ideation, with potentially important implications for developer productivity and participation i…