2 papers
cs.SE2026
RubberDuckBench: A Benchmark for AI Coding Assistants
Ferida Mohammed, Fatma Ayad, Petros Maniatis +2
Programmers are turning to AI coding assistants to answer questions about their code. Benchmarks are needed to soundly evaluate these systems and understand their performance. To e…
cs.SE2024
CRQBench: A Benchmark of Code Reasoning Questions
Elizabeth Dinella, Satish Chandra, Petros Maniatis
Large Language Models have demonstrated exceptional proficiency on coding tasks, but it is challenging to precisely evaluate their code reasoning ability. Existing benchmarks are i…