3 papers
cs.SE2026
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
Gabriel Orlanski, Devjeet Roy, Alexander Yun +7
Software development is iterative, yet agentic coding benchmarks hide design issues through their single-shot setup. Recent iterative benchmarks attempt to remedy this but heavily…
cs.SE2025
Context-Specific Instruction: A Longitudinal Study on Debugging Skill Acquisition and Retention for Novice Programmers
Ziyi Zhang, Devjeet Roy, Venera Arnaoudova
Bug localization is a critical skill, yet novices often lack systematic approaches. Prior work tested abstract guidelines and general concrete steps; the impact of context-specific…
cs.LG2025
Conformal Prediction Sets for Deep Generative Models via Reduction to Conformal Regression
Hooman Shahrokhi, Devjeet Raj Roy, Yan Yan +2
We consider the problem of generating valid and small prediction sets by sampling outputs (e.g., software code and natural language text) from a black-box deep generative model for…