3 papers
cs.SE2026
RubberDuckBench: A Benchmark for AI Coding Assistants
Ferida Mohammed, Fatma Ayad, Petros Maniatis +2
Programmers are turning to AI coding assistants to answer questions about their code. Benchmarks are needed to soundly evaluate these systems and understand their performance. To e…
cs.SE2026
Wink: Recovering from Misbehaviors in Coding Agents
Rahul Nanda, Chandra Maddila, Smriti Jha +3
Autonomous coding agents, powered by large language models (LLMs), are increasingly being adopted in the software industry to automate complex engineering tasks. However, these age…
cs.SE2025
ECO: An LLM-Driven Efficient Code Optimizer for Warehouse Scale Computers
Hannah Lin, Martin Maas, Maximilian Roquemore +14
With the end of Moore's Law, optimizing code for performance has become paramount for meeting ever-increasing compute demands, particularly in hyperscale data centers where even sm…