collaborators

8 papers

cs.HC2026

Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents

Aayush Kumar, Avik Dutta, Sumit Gulwani +3

Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the agent to develop a plan before task execu…

cs.SE2026

Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI

Emerson Murphy-Hill, Jenna Butler, Alexandra Savelieva

Organizations rolling out agentic command line tools like Anthropic's Claude Code and GitHub's Copilot CLI need to know who will try them, who will keep using them, and whether the…

cs.SE2026

GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis

Alex Heilman, Alex Kyllo, Emerson Murphy-Hill

Does GitHub Copilot (GHCP) make engineers more productive, or do the engineers who use it more differ from those who use it less? And even within a single engineer, are GHCP-heavy…

cs.SE2026

Future of Software Engineering Research: The SIGSOFT Perspective

Massimiliano Di Penta, Kelly Blincoe, Marsha Chechik +4

As software engineering conferences grow in size, rising costs and outdated formats are creating barriers to participation for many researchers. These barriers threaten the inclusi…

cs.SE2025

SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks

Sanket Mhatre, Yasharth Bajpai, Sumit Gulwani +2

AI coding agents have shown great progress on Python software engineering benchmarks like SWE-Bench, and for other languages like Java and C in benchmarks like Multi-SWE-Bench. How…

cs.SE2025

Why AI Agents Still Need You: Findings from Developer-Agent Collaborations in the Wild

Aayush Kumar, Yasharth Bajpai, Sumit Gulwani +2

Software Engineering Agents (SWE agents) can autonomously perform development tasks on benchmarks like SWE Bench, but still face challenges when tackling complex and ambiguous real…