4 papers
From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale
Chandra Maddila, Mashrur Rashik, Euna Mehnaz Khan +4
AI coding agents are generating code at volumes that exceed the capacity of traditional peer review. At the same time, existing AI code review tools over-index on low-value suggest…
REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
Smriti Jha, Matteo Paltenghi, Chandra Maddila +3
Production deployment of AI coding agents requires fast, reproducible evaluation signals. Existing industrial practices trade off speed and fidelity: online A/B testing takes weeks…
Developing and evaluating a chatbot to support maternal health care
Smriti Jha, Vidhi Jain, Jianyu Xu +8
The ability to provide trustworthy maternal health information using phone-based chatbots can have a significant impact, particularly in low-resource settings where users have low…
Wink: Recovering from Misbehaviors in Coding Agents
Rahul Nanda, Chandra Maddila, Smriti Jha +3
Autonomous coding agents, powered by large language models (LLMs), are increasingly being adopted in the software industry to automate complex engineering tasks. However, these age…