10 citations · 14 across the 4 of their papers we have counts for
4 papers
From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering
Tao Dong, Harini Sampath, Ja Young Lee +2
As Large Language Models (LLMs) evolve from code generators into collaborative partners for software engineers, our methods for evaluation are lagging. Current benchmarks, focused…
What do professional software developers need to know to succeed in an age of Artificial Intelligence?
Matthew Kam, Cody Miller, Miaoxin Wang +8
Generative AI is showing early evidence of productivity gains for software developers, but concerns persist regarding workforce disruption and deskilling. We describe our research…
Creating benchmarkable components to measure the quality ofAI-enhanced developer tools
Elise Paradis, Ambar Murillo, Maulishree Pandey +4
In the AI community, benchmarks to evaluate model quality are well established, but an equivalent approach to benchmarking products built upon generative AI models is still missing…
How much does AI impact development speed? An enterprise-based randomized controlled trial
Elise Paradis, Kate Grey, Quinn Madison +6
How much does AI assistance impact developer productivity? To date, the software engineering literature has provided a range of answers, targeting a diversity of outcomes: from per…