5 papers
The Perfection Paradox: From Architect to Curator in AI-Assisted API Design
Mak Ahmad, Andrew Macvean, JJ Geewax +1
Enterprise API design is often bottlenecked by the tension between rapid feature delivery and the rigorous maintenance of usability standards. We present an industrial case study e…
From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering
Tao Dong, Harini Sampath, Ja Young Lee +2
As Large Language Models (LLMs) evolve from code generators into collaborative partners for software engineers, our methods for evaluation are lagging. Current benchmarks, focused…
What do professional software developers need to know to succeed in an age of Artificial Intelligence?
Matthew Kam, Cody Miller, Miaoxin Wang +8
Generative AI is showing early evidence of productivity gains for software developers, but concerns persist regarding workforce disruption and deskilling. We describe our research…
Creating benchmarkable components to measure the quality ofAI-enhanced developer tools
Elise Paradis, Ambar Murillo, Maulishree Pandey +4
In the AI community, benchmarks to evaluate model quality are well established, but an equivalent approach to benchmarking products built upon generative AI models is still missing…
How much does AI impact development speed? An enterprise-based randomized controlled trial
Elise Paradis, Kate Grey, Quinn Madison +6
How much does AI assistance impact developer productivity? To date, the software engineering literature has provided a range of answers, targeting a diversity of outcomes: from per…