4 papers · 1 filter
The Perfection Paradox: From Architect to Curator in AI-Assisted API Design
Mak Ahmad, Andrew Macvean, JJ Geewax +1
Enterprise API design is often bottlenecked by the tension between rapid feature delivery and the rigorous maintenance of usability standards. We present an industrial case study e…
From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering
Tao Dong, Harini Sampath, Ja Young Lee +2
As Large Language Models (LLMs) evolve from code generators into collaborative partners for software engineers, our methods for evaluation are lagging. Current benchmarks, focused…
Creating benchmarkable components to measure the quality ofAI-enhanced developer tools
Elise Paradis, Ambar Murillo, Maulishree Pandey +4
In the AI community, benchmarks to evaluate model quality are well established, but an equivalent approach to benchmarking products built upon generative AI models is still missing…
How much does AI impact development speed? An enterprise-based randomized controlled trial
Elise Paradis, Kate Grey, Quinn Madison +6
How much does AI assistance impact developer productivity? To date, the software engineering literature has provided a range of answers, targeting a diversity of outcomes: from per…